时刻更智能:环境驱动的动态策略实现大语言模型持续改进
Smarter by the Moment: Environment-Driven Dynamic Policies for Continual LLM Improvement
浏览论文内容
中文总结 AI 辅助
针对大语言模型持续适应难题,提出动态检索策略生成框架(DRPG),利用历史数据与环境反馈生成任务特定策略,在六个基准上超越强基线,并揭示策略指导有效性的条件。
中文摘要 AI 辅助
大语言模型(LLMs)在多个领域取得了显著进展,但持续适应不断演变的任务和环境仍是一个关键挑战。现有的记忆增强方法将单个历史示例作为直接参考进行检索,但并未从这些示例中明确综合出可操作的策略,导致相同类型的错误反复出现。我们提出了动态检索策略生成(DRPG)框架,该框架将基于记忆的检索与动态策略生成器相结合,利用历史数据和环境反馈为持续的大语言模型改进生成特定于任务的策略。我们在涵盖文本到SQL、问答、医学诊断和Python编程的六个基准上,使用来自专有和开源权重家族的七种大语言模型评估了DRPG。DRPG在大多数数据集和模型上优于强基线。进一步分析表明,DRPG的策略生成对检索策略具有鲁棒性,无需先前的策略连续性即可有效运行,并且可以利用较小或跨家族的模型作为成本效益高的策略生成器。我们还发现,策略级指导的益处取决于任务特征,这为该方法在何时以及何种条件下最有效提供了实用见解。
英文摘要
Large Language Models (LLMs) have achieved remarkable progress across diverse domains, but continual adaptation to evolving tasks and environments remains a key challenge. Existing memory-augmented approaches retrieve individual past examples as direct references, but do not explicitly synthesize actionable strategies from them, causing the same types of errors to recur. We propose Dynamic Retrieval-based Policy Generation (DRPG), a framework that integrates memory-based retrieval with a dynamic policy generator, leveraging historical data and environment feedback to produce task-specific policies for continual LLM improvement. We evaluate DRPG across six benchmarks spanning text-to-SQL, question answering, medical diagnosis, and Python programming, using seven LLMs from both proprietary and open-weight families. DRPG outperforms strong baselines across most datasets and models. Further analysis demonstrates that DRPG's policy generation is robust to retrieval strategy, operates effectively without prior policy continuity, and can leverage smaller or cross-family models as cost-efficient policy generators. We also find that the benefit of policy-level guidance depends on task characteristics, offering practical insights into when and under what conditions this mechanism is most effective.
发表机构
- National Taiwan University(国立台湾大学)
- Academia Sinica(中央研究院)
- AI Research Center (AINTU), National Taiwan University(国立台湾大学人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。