超越脚本化搜索:通过智能体黑盒优化实现样本高效的奖励发现
Beyond Scripted Search: Sample-Efficient Reward Discovery via Agentic Black-box Optimization
浏览论文内容
中文总结 AI 辅助
针对强化学习奖励搜索样本效率低的问题,提出ARBO框架,利用LLM智能体基于评估历史动态构建搜索策略,在四个控制领域显著提升性能。
中文摘要 AI 辅助
为低层强化学习(RL)控制设计稠密奖励函数仍然困难。近期工作利用大型语言模型(LLM)在脚本化搜索算法中,根据策略训练反馈迭代生成并改进奖励函数。然而,评估每个候选方案需要一次完整的RL训练运行,这使得样本效率成为复杂控制任务上奖励搜索的核心挑战。为解决此局限,我们提出智能体奖励黑盒优化(ARBO)框架,其中LLM智能体在运行时根据评估历史构建搜索策略,该历史作为其持久工作区维护。评估历史包含两个组成部分:由评估预言机维护的观测,包括候选分数、逐项训练曲线和错误回溯;以及智能体维护的信念,记录诊断和预期下一步。智能体通过工具查询两者并生成下一批奖励候选,而非从固定提示中一次性生成。在四个控制领域中,在共享评估预算下,ARBO在操作成功率上获得29.9%的提升,在电网得分上获得192.8%的提升。消融研究考察了每个组件的贡献及对骨干选择的敏感性。
英文摘要
Designing dense reward functions for low-level reinforcement learning (RL) control remains difficult. Recent work uses large language models (LLMs) to iteratively generate and refine reward functions using policy-training feedback within scripted search algorithms. However, evaluating each candidate requires a full RL training run, making sample efficiency a central challenge for reward search on complex control tasks. To address this limitation, we propose an Agentic Reward Black-box Optimization (ARBO) framework, in which an LLM agent builds the search strategy at run time from an evaluation history maintained as its persistent workspace. The evaluation history comprises two components: observations maintained by the evaluation oracle, including candidate scores, per-term training curves, and error tracebacks; and an agent-maintained belief that records diagnoses and intended next steps. The agent queries both with tools and generates the next batch of reward candidates, rather than generating them in a single pass from a fixed prompt. Across four control domains, ARBO achieves gains of 29.9% in manipulation success rate and 192.8% in power-grid score over baseline means under a shared evaluation budget. Ablations examine each component's contribution and sensitivity to backbone choice.