Mara Chain:将失败重新构想为AI系统自动进化的垫脚石
Mara Chain: Rethinking Failure as a Stepping Stone for AI System Auto-Evolution
- Ant Group(蚂蚁集团)
- Beijing Intelligent Game and Decision Lab(北京智能博弈与决策实验室)
- Beijing Defense Innovation Institute(北京国防创新研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
Mara Chain通过保留并迭代细化被拒绝的候选配置,将失败转化为优化垫脚石,在多个任务上以更少回滚实现更高性能提升。
AI中文摘要:
优化已部署的AI系统越来越等同于编辑提示、技能、工具链和代码,而非模型权重。现有方法通常通过“提出-评估-选择”流程来优化这些工件,其中候选配置被评估,只有满足接受标准的配置才会被选中。然而,我们的分析表明,被丢弃的候选往往包含对后续优化至关重要的信息。丢弃它们会导致后续提案重复遇到相同的失败模式。我们引入了Mara Chain,一种将拒绝的候选转化为垫脚石的细化流程。Mara Chain不是丢弃被拒绝的候选,而是保留它,并利用先前尝试中积累的证据对其进行迭代细化。该流程将每个细化链限制在固定深度,并应用帕累托过滤的Top-N选择来限制候选池。在AppWorld技能优化、TerminalBench 2.1工具链优化和MuSiQue检索流水线优化中,Mara Chain以更少的回滚次数带来了更大的任务性能提升。它在AppWorld上相对性能比GEPA、ACE和SkillOpt-Lite高出最多20.5%,达到目标分数所需的回滚次数比GEPA少65.5%。在TerminalBench 2.1上,它相对于AHE和Meta-Harness分别将通过率提高了20.2和22.5个百分点,并在MuSiQue测试中将nDCG@10和Recall@10分别提高了0.104和0.131,优于手写检索流水线。
英文摘要:
Optimizing deployed AI systems increasingly amounts to editing prompts, skills, harnesses, and code rather than model weights. Existing approaches commonly optimize these artifacts through propose-evaluate-select procedures, where candidate configurations are evaluated and only those meeting an acceptance criterion are selected. Yet our analysis shows that discarded candidates often contain information critical for subsequent optimization. Discarding them causes later proposals to revisit the same failure modes. We introduce Mara Chain, a refinement procedure that turns rejected candidates into stepping stones. Rather than discarding a rejected candidate, Mara Chain retains and iteratively refines it using evidence accumulated across preceding attempts. The procedure limits each refinement chain to a fixed depth and applies Pareto-filtered Top-N selection to bound the candidate pool. Across AppWorld skill optimization, TerminalBench 2.1 harness optimization, and MuSiQue retrieval-pipeline optimization, Mara Chain delivers greater task-performance gains with fewer rollouts. It outperforms GEPA, ACE, and SkillOpt-Lite by up to 20.5% in relative performance on AppWorld, reaching the target score with 65.5% fewer rollouts than GEPA. It improves the pass rate by 20.2 and 22.5 percentage points over AHE and Meta-Harness on TerminalBench 2.1, respectively, and improves MuSiQue test nDCG@10 and Recall@10 by 0.104 and 0.131 over a hand-written retrieval pipeline.