arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LongWoF-Bench:评估用于可验证长工作流任务的EvoMap基因

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

Xiao Zhang, Qumeng Sun, Jiahao Li, Xiang Liu, Bruno, Haoyang Zhang

arXiv 2608.23200首次发表:更新:

发表机构

EvoMap; Tsinghua University(EvoMap; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出LongWoF-Bench基准,发现通过EvoMap将验证的执行轨迹转化为可复用的Gene,能显著提升多模型在长工作流任务上的表现,且优于Skill和参考蒸馏的Gene。

AI 中文摘要

大型语言模型日益被期望执行复杂工作流,其成功取决于维持相互依赖的约束并生成满足严格端到端验证的产物。然而,成功的执行经验通常在单次运行后就会丢失,迫使后续模型从头重新发现策略和失败模式。我们研究这种经验是否可以通过EvoMap进行外部化和复用,其中验证器确认的执行轨迹被整合为结构化的基因(Gene)。为评估该设置,我们引入了长工作流基准(Long-Workflow Benchmark,LongWoF-Bench),涵盖代码生成、智能体-环境合成、数学推理和规则遵循领域的778项可机器验证任务。在具有验证器确认的Opus轨迹的252项任务上,进化后的EvoMap基因在所有7个评估模型上的表现均优于技能(Skill),优势幅度为8.7至15.5个百分点,且该优势延伸至不同模型家族的消费级模型。相比之下,通过参考蒸馏得到的基因并未表现出相同的优势,这表明仅靠紧凑表示是不够的,基因的效用与验证经验的来源密切相关。对于Claude Opus,复用基因还比技能多完成39项任务,同时将求解时间的令牌消耗降低了9.9%。综合来看,这些结果表明,经过验证的执行经验可以作为可复用的外部资源被保留和共享,使模型能够在不重复支付经验发现的全部成本的情况下提升长工作流的完成能力。

英文摘要

Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.

CommentsTechnical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑