arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22086cs.AIcs.CV

Designer-RSI:从用户流量中演化程序性记忆用于智能体图形设计

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

  • Adobe
  • Brown University(布朗大学)

机构由 AI 辅助整理,请以论文原文为准。

Hongyang Du, Lan Yan, Christian Flores, Asim Kadav

AI总结:

提出Designer-RSI框架,通过外部程序性记忆从用户流量中持续演化设计技能,在无权重更新和人工标签下,将GenEval2成功率从72.7%提升至99.3%,并显著提高多个基准的胜率。

AI中文摘要:

专业图形设计是一项长视界的智能体任务,其中结构化、可编辑的工件由许多相互依赖的操作涌现,然而结果没有可靠的可编程评估器。我们引入了一个持续适应框架,其中冻结的前沿模型通过超过230个工具操作专业设计软件,而外部程序性记忆存储自然语言技能,从经验中积累并完善可复用的设计程序。该记忆通过获取重复出现的未覆盖子任务的程序而拓宽,并通过根据自身成功和失败执行修订现有程序而深化,同时匹配的重放门仅允许修复失败而不回归观察到的成功的更改。在1,406个真实用户简报和1,869个自动评分轨迹上进行了五轮,无权重更新且无人工标签,将技能库从76个文档衍生技能增长到139个,并将Claude-Sonnet-4上的GenEval2执行成功率从72.7%提升到99.3%(生成质量提高+11.99分),在四个专业设计基准上对Claude-Sonnet-4和Claude-Opus-4.6分别达到61.8%和67.6%的胜率,对比无技能智能体。我们进一步展示了两种机制组合的有效性:在用户流量基准的200个保留简报上,仅拓宽或仅深化分别达到49.4%和48.6%的胜率,对比无技能智能体,而组合达到58.5%(p=0.025)。程序性记忆为智能体在嘈杂、不可验证反馈下的持续适应提供了一条实用途径。

英文摘要:

Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring procedures for recurring uncovered subtasks and deepens by revising existing procedures against their own successful and failed executions, while a matched replay gate admits only changes that repair failures without regressing observed successes. Five rounds over 1,406 real user briefs and 1,869 automatically graded trajectories, with no weight updates and no human labels, grow the bank from 76 documentation-derived skills to 139 and raise GenEval2 execution success on Claude-Sonnet-4 from 72.7% to 99.3% (+11.99 points in generation quality), with 61.8% and 67.6% win rates against the no-skill agent across four specialized design benchmarks on Claude-Sonnet-4 and Claude-Opus-4.6. We further show the two mechanisms are effective in combination: on 200 held-out briefs from user-traffic benchmark, widening or deepening alone reaches a 49.4% / 48.6% win rate over the no-skill agent, while their combination reaches 58.5% (p = 0.025). Procedural memory offers a practical route to continual adaptation of agents under noisy, unverifiable feedback.

补充信息

相关深度报道

↑