arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EngramBench:一个基于能力锚定的技能进化框架基准

EngramBench: A Capability-Grounded Benchmark for Skill-Evolution Harnesses

Zhixuan Tan, Pengjie Gu, Zhao Li, Yihan Hu, Xu He, Dong Li, Jianye Hao

arXiv 2609.39284首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; MemoraX AI(香港中文大学(深圳); MemoraX人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体技能进化评估中能力抽象与解决方案泄漏混淆的问题,提出EngramBench基准,通过30个学习任务和13个迁移任务验证,证明静态技能库虽不能突破推理瓶颈,但能减少55%以上编码时间,实现跨领域能力迁移的可验证测量。

AI 中文摘要

尽管大型语言模型在孤立代码生成方面取得了显著成功,但真实的软件工程需要持续的推理、复杂的状态管理和跨领域的连续抽象。然而,当前对自主智能体技能进化的评估存在一个关键的识别性问题:它们在结构上将真正的能力抽象与机械式的解决方案泄漏(即从历史训练数据中复制高度相似的代码)混为一谈。为解决这一问题,我们引入了EngramBench,这是一个严格且基于能力锚定的基准,遵循“能力重叠但解决方案不重叠”的严格公理。EngramBench包含30个多样的学习任务和13个未见过的迁移任务,挑战智能体在由LLM模拟用户驱动的交互式、多小时的开发周期中进行导航。我们广泛的评估覆盖了48条多小时的执行轨迹,并经人类专家验证,揭示了关于程序性记忆的深刻见解。我们证明,静态技能库并不能神奇地绕过精确代码实现的“最后一公里”,这一环节仍受限于基础模型固有的推理能力瓶颈。然而,它们作为不可或缺的执行指南针,通过引导智能体远离灾难性的、高令牌消耗的试错过程,真正的能力抽象削减了冗余的上下文膨胀,并将整体编码时间减少了超过55%。最终,EngramBench将评估范式从琐碎的模式匹配转变为对深层跨领域能力迁移的可验证测量。

英文摘要

While large language models have achieved remarkable success in isolated code generation, authentic software engineering requires sustained reasoning, complex state management, and continuous cross-domain abstraction. However, current evaluations of skill evolution in autonomous agents suffer from a critical identifiability problem: they structurally confound genuine capability abstraction with rote solution leakage (i.e., copying highly similar code from historical training data). To resolve this, we introduce EngramBench, a rigorous, capability-grounded benchmark governed by the strict axiom of capability overlap without solution overlap. Comprising 30 diverse learning tasks and 13 unseen transfer tasks, EngramBench challenges agents to navigate interactive, multi-hour development cycles driven by LLM-simulated users. Our extensive evaluation across 48 multi-hour execution trajectories -- corroborated by human-expert validation -- reveals a profound insight into procedural memory. We demonstrate that static skill banks do not magically bypass the "last mile" of exact code implementation, which remains bottlenecked by the base model's inherent reasoning limits. However, they serve as an indispensable execution compass. By navigating agents away from catastrophic, token-heavy trial-and-error, genuine capability abstraction slashes redundant context bloat and reduces overall coding time by over 55%. Ultimately, EngramBench shifts the evaluation paradigm from trivial pattern matching to the verifiable measurement of deep, cross-domain capability transfer.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑