NLPG:用于自进化语言智能体的自然语言策略梯度
NLPG: Natural-Language Policy Gradients for Self-Evolving Language Agents
浏览论文内容
中文总结 AI 辅助
提出自然语言策略梯度(NLPG),一种外部策略记忆方法,通过诊断执行轨迹并将失败转化为局部自然语言修正,在不改变参数或结构的情况下改进冻结的智能体,在六个基准上平均超越最强基线8.71个百分点。
中文摘要 AI 辅助
大型语言模型智能体越来越依赖于复合程序来进行检索、工具使用、推理和验证,然而它们的失败往往源于局部程序性决策。现有的强化学习和提示优化方法通常依赖于标量奖励或反复修改整个提示,这使得在保持智能体冻结的同时难以捕获和重用程序性改进。为解决此问题,我们提出了自然语言策略梯度(NLPG),一种外部策略记忆方法,用于在不改变模型参数或程序结构的情况下改进固定智能体。NLPG诊断执行轨迹,通过模块图向后传播下游反馈,并将反复出现的失败转化为路由局部的自然语言修正,这些修正被聚合成有界的策略更新,用于后续执行。在涵盖记忆、推理、指令遵循和证据验证的六个基准测试中,NLPG还平均比每个基准的最强列出的基线高出8.71个百分点。这些结果提供了证据,表明评估过的程序性经验可以被转化为局部且可解释的策略更新,从而实现冻结智能体的持续改进。
英文摘要
Large language model agents increasingly rely on compound programs for retrieval, tool use, reasoning, and verification, yet their failures often arise from local procedural decisions. Existing reinforcement-learning and prompt-optimization approaches typically rely on scalar rewards or repeatedly modify entire prompts, making it difficult to capture and reuse procedural improvements while preserving a frozen agent. To address this problem, We propose Natural-Language Policy Gradients (NLPG), an external policy-memory method for improving a fixed agent without changing its model parameters or program structure. NLPG diagnoses execution traces, propagates downstream feedback backward through the module graph, and converts recurring failures into route-local natural-language corrections that are aggregated into bounded policy updates for subsequent executions. Across six benchmarks covering memory, reasoning, instruction following, and evidence verification, NLPG also outperforms the strongest listed baseline for each benchmark by 8.71 percentage points on average. These results provide evidence that evaluated procedural experience can be transformed into local and interpretable policy updates, enabling continual improvement of frozen agents.
发表机构
- Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。