RLMOpt:基于递归语言模型的自适应提示优化器
RLMOpt: Adaptive Prompt Optimization via Recursive Language Models
- Autonomize AI
- Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出基于递归语言模型(RLM)的自适应提示优化器RLMOpt,在四个基准上相比GEPA表现更优,效率更高且提示更小。
AI中文摘要:
提示优化器可自动搜索能提升语言模型性能的提示,但现有方法依赖预定义的优化流程:算法决定探索哪些候选、搜索如何推进,而语言模型生成或优化提示提案。本文提出RLMOpt,一种通过递归语言模型(RLM)使搜索策略本身由语言模型驱动的提示优化器。RLM智能体在基于工具的环境中运行,检查任务信息、分析失败情况、生成候选、分配评估预算并决定何时停止;确定性控制模块补充智能体,强制执行客观评分、基于帕累托的选择和回归约束。我们在四个基准上评估RLMOpt,涵盖结构化临床信息提取(Chia)、多跳问答(HotpotQA)、可验证指令遵循(IFBench-2025)和多轮工具调用智能体(BFCL)。在单种子匹配比较中,RLMOpt在所有四个基准上取得最佳保留分数,四项任务均值为0.610,而GEPA为0.589。跨种子重复每个基准共得到11次匹配的基准-种子比较,其中RLMOpt在9次案例中优于GEPA;在全部11次运行中,RLMOpt从未产生性能低于其种子的提示,而GEPA有两次低于其起点。RLMOpt效率更高,用更少的搜索rollout取得这些结果,同时生成的提示大小为GEPA的27%-79%。我们的结果进一步表明,优化增益主要由种子提示中可用的提升空间决定,而非搜索预算,因此高效优化依赖于可靠地达到可用提升空间且搜索量最小。
英文摘要:
Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization procedure: the algorithm determines which candidates to explore and how the search progresses, while the language model generates or refines prompt proposals. We introduce RLMOpt, a prompt optimizer that makes the search policy itself language-model-driven through a recursive language model (RLM). The RLM agent operates over a tool-based environment, inspecting task information, analyzing failures, generating candidates, allocating evaluation budget, and deciding when to stop. A deterministic harness complements the agent by enforcing objective scoring, Pareto-based selection, and regression constraints. We evaluate RLMOpt across four benchmarks spanning structured clinical information extraction (Chia), multi-hop question answering (HotpotQA), verifiable instruction following (IFBench-2025), and multi-turn tool-calling agents (BFCL). In a matched comparison at a single seed, RLMOpt obtains the best held-out score on all four benchmarks and leads the four-task mean (0.610 against 0.589 for GEPA). Repeating each benchmark across seeds yields 11 matched benchmark-seed comparisons, in which RLMOpt outperforms GEPA in 9 cases. Across all 11 runs, it never produced a prompt that underperformed its seed, whereas GEPA fell below its starting point twice. It is also more efficient, achieving these results with fewer search rollouts while producing prompts that are 27-79% the size of those produced by GEPA. Our results further show that optimization gains are determined primarily by the headroom available in the seed prompt, rather than by the search budget. Efficient optimization therefore depends on reaching the available headroom reliably and with minimal search