发表机构
Alibaba Cloud Computing(阿里云)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究推出面向金融工作流的FinEvo-Bench基准,对比4种自进化智能体框架,发现进化可提升性能、减少合规问题,且评分规则反馈效果优于参考答案反馈。
AI 中文摘要
大多数智能体基准独立评估任务,无法衡量一项任务的经验是否有助于后续任务。现有的自进化基准未同时覆盖专业工作流、开放式交付物及多维度评估。我们推出FinEvo-Bench,这是一个纵向基准,包含120个基于真实案例的任务,覆盖6个金融领域的20个业务场景。任务所需操作和约束由机构提供的专业流程定义,合格的机构提供及公开记录的案例提供任务事实。每个场景包含6个相关但实质不同的案例,共享同一专业流程及人工审核的任务质量与金融合规性评分规则。我们使用相同的Qwen3.7-Max主干模型,搭配3组独立打乱、全局交错的任务流,对比4种自进化智能体框架。配对的非进化对照组用于评估每个框架从保留经验中获得的自进化增益,由Claude Opus 4.6支持的独立Claude Code评分智能体评估所有输出。Letta获得最高的进化后分数(91.65)及最少的合规问题(每个任务0.09个);Codex实现最大的自进化增益(+19.37)。在所有框架中,进化条件使分数提升9.33-19.37分,合规问题减少每个任务0.12-0.44个。在场景内排名4-6的配对分数增益比排名1-3的高6.10-8.70分。在Claude Code中,仅技能进化比仅记忆进化及记忆-技能结合进化产生更高的任务质量和更少的合规问题。在所有4种框架中,评分规则反馈比参考答案反馈产生更高的分数和更少的合规问题。FinEvo-Bench同时衡量专业性能和自进化能力,即智能体将先前经验转化为后续改进的效率。
英文摘要
Agents used over time encounter recurring professional work: each case requires different evidence and judgment, while the underlying workflow can be reused. Benchmarks built from independent tasks cannot reveal whether an agent turns earlier experience into better procedures for later cases. We introduce FinEvo-Bench, a longitudinal benchmark designed around this structure. It contains 120 open-ended tasks drawn from real cases across 20 business scenes in six financial domains. Each scene contains six substantively different cases that share a professional workflow and an expert-authored rubric for task quality and financial compliance. Constructing and validating the benchmark required approximately 1,200 person-hours. Finance provides a natural test bed because recurring analyses apply shared professional and compliance requirements to heterogeneous inputs, producing case-specific analyses and conclusions. We evaluate four self-evolving agent scaffolds with Qwen3.7-Max on three independently shuffled, globally interleaved task streams. A Claude Code rubric judge backed by Claude Opus~4.6 evaluates all outputs, and paired state-reset controls estimate each scaffold's gain from retained experience. Evolving runs score 9.33--19.37 points higher and trigger 0.12--0.44 fewer compliance issues per task than their paired controls. Paired score gains at within-scene ranks~4--6 exceed those at ranks~1--3 by 6.10--8.70 points. FinEvo-Bench measures whether retained experience improves later professional work under continued use.
Comments22 pages, 4 figures; includes appendices