arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.09044cs.CL

经验树:面向自我进化智能体的分层经验管理

Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents

Zihao Deng, Yining Zhu, Leiming Wang, Junbo Wang, Jingfei Lu

AI总结:

该研究提出经验树(ToE)框架,将经验组织与LLM智能体分层推理对齐,在“24点游戏”和“金融进化基准”上大幅提升了解题性能与效率,解决了现有经验表示与推理过程脱节的问题。

AI中文摘要:

持续自我进化要求大语言模型(LLM)智能体将环境交互转化为可靠且可复用的经验。现有方法通常优化单个轨迹或从相关轨迹中抽象共享知识,但它们的经验表示往往与底层推理过程脱节,这限制了反馈归因、跨任务迁移以及更新与检索效率,尤其在仅具备结果级反馈的复杂推理任务中。为克服这一局限,我们提出**经验树(Tree-of-Experience, ToE)**,一种结构化的经验管理框架,其将经验组织与LLM智能体的分层推理过程对齐。具体而言,ToE将经验组织为一个由分析视角与推理路径构成的共享树,其可靠性通过环境结果校准,以支持系统化更新、迁移与高效检索。在“24点游戏(Game of 24)”和“金融进化基准(FinEvolveBench)”上的实验结果表明,ToE显著提升了解题性能与效率:在“24点游戏”上,ToE相较于无经验的思维树(ToT)基线实现了31.4%的准确率相对提升;在“金融进化基准”上,ToE在12个评估设置下相较于无经验流水线平均提升了41.24%的tsIC,而传统经验管理方法的表现往往不及无经验基线。

英文摘要:

Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectories or abstract shared knowledge from related trajectories, but their experience representations are often disconnected from the underlying reasoning process. This limits feedback attribution, cross-task transfer, and update and retrieval efficiency, particularly in complex reasoning tasks with outcome-level feedback. To overcome this limitation, we propose \textbf{T}ree-\textbf{o}f-\textbf{E}xperience (ToE), a structured experience-management framework that aligns experience organization with the hierarchical reasoning process of LLM agents. Specifically, ToE organizes the experience into a shared tree of analytical perspectives and reasoning paths, whose reliability is calibrated through environmental outcomes to support systematic updating, transfer, and efficient retrieval. The experimental results on \textsc{Game of 24} and \textsc{FinEvolveBench} show that ToE substantially improves both problem-solving performance and efficiency. On \textsc{Game of 24}, ToE achieves a 31.4\% relative improvement in accuracy over the experience-free ToT baseline. On \textsc{FinEvolveBench}, ToE improves tsIC by an average of 41.24\% over the experience-free pipeline across 12 evaluation settings, whereas conventional experience-management methods often underperform experience-free baselines.

↑