AhaBench:智能体能否从先前经验中学习?一个面向长时程持续学习的基准
AhaBench: Do Agents Turn Experience into Reusable Insights? A Long-Horizon Benchmark for Continual Learning
另 2 家 · 查看机构详情
- Princeton University(普林斯顿大学)
- Tencent Hy(腾讯Hy)
- Tsinghua University(清华大学)
- Shanghai Jiaotong University(上海交通大学)
- Hong Kong University(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
AhaBench提出长时程持续学习基准,通过三个任务测试智能体在支持移除或延迟后能否利用经验提升表现,并引入初始分数、经验后分数和学习提升三部分评分,揭示模型能力差异。
中文摘要 AI 辅助
现代语言智能体被期望在长时程中运行:它们会提出后续问题、复用工作示例、处理工具反馈,并适应延迟的后果。大多数评估仍然在提示后重置智能体,或仅对单个轨迹的最终状态进行评分。AhaBench提出了一个更具操作性的问题:当固定模型获得有用经验时,在相关评估条件下(明显支持已被移除、更改或延迟),其后续行为是否会改善?该套件包含三个组成部分。Aha-Puzzle测试在解决隐藏状态谜题后的无提示探索;Aha-Euler将Project-Euler风格的数学思想转化为生成的教授/保留任务,并配备精确验证器;Aha-Vending是一个受Vending-Bench启发的开源实现,测试模拟售货智能体在处理延迟反馈和运营事件时是否保持盈利。AhaBench报告三部分记分卡:初始分数衡量起始能力,经验后分数衡量后续经验结果,学习提升是两者的差值。这种分解是主要的实证信息:善于利用可见支持的模型、达到高经验后分数的模型以及在运行中提升最多的模型并不总是相同的。在常见的八模型面板上,Claude Opus 4.6以64.3的总体经验后分数和+25.8的总体学习提升领先,Gemini 3.1 Pro以63.4紧随其后。各组成部分的结果解释了这种差异:谜题轨迹提高了有支持分数,但往往未能转化为无提示探索行为;Aha-Euler完整教学达到78.6-100.0%,而仅答案迁移范围为0.0-73.9%;Aha-Vending将盈利的事件处理与破产和无订单失败区分开来。我们发布了基准任务、评分标准、验证器、模拟器代码以及用于评估新智能体的接口。
英文摘要
Can language agents continually learn from experience, turning earlier interactions into reusable capabilities? AhaBench evaluates this ability through exploration after solved hidden-state puzzles, computational transfer after mathematical teaching, and sustained business operation under delayed feedback. The benchmark is agnostic to how an agent learns; the evaluated agents use fixed model weights. Curriculum profiles, teaching contrasts, and daily trajectories reveal a common challenge: using explicit guidance is more reliable than generalizing beyond it or sustaining useful behavior. Across the Puzzle panel, the advantage over matched cold targets is 36.0-53.5 points greater with trace support than at the trace-free endpoint; Qwen 3.6 Plus nevertheless retains a +12.57-point post-curriculum gain. In Euler, worked procedures yield 80.0-100.0% held-out accuracy across models, while question-plus-answer teaching yields 0.0-73.9%. Vending trajectories separate sustained profit, late recovery, and incomplete operation: Doubao Seed 2.0 Pro finishes nominal operation at +495 but averages -10 over the year. Together, these results make continual learning an operational target: experience should yield capabilities that remain effective as guidance, inputs, and business states change. We release tasks, validators, a simulator, records, and analyses for developing agents that turn useful insights into lasting abilities.