arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于使冻结的语言模型智能体学习领域的控制系统、数据集和方法

A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain

Debjyoti Paul

arXiv 2607.25415首次发表:更新:

发表机构

Meta AI(Meta AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何让冻结的语言模型智能体学习领域,核心方法是将框架视为特定动作空间,用经典强化学习在线学习策略并以多目标奖励评分,贡献包括用DSPy实例化系统并评估,还发布相关代码、任务套件、训练日志及部署方法。

AI 中文摘要

生产语言模型智能体越来越多地由包装在框架中的冻结模型组装而成,包括提示模板、工具集等。2026年的两个系统Meta-Harness和HyperAgents表明框架可被智能提议者优化或重写,但存在代价。我们采取更受限立场,将框架视为小的、固定的、人类可读的动作空间,用经典样本高效强化学习在线学习策略,以多目标奖励评分。我们用DSPy实例化该控制系统并在三个可验证任务领域和两个模型提供者上评估,还发布了相关代码、任务套件、训练日志和部署方法。

英文摘要

Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieval layer, a planning strategy, and a verification policy. Two 2026 systems, Meta-Harness (Lee et al., 2026) and HyperAgents (Meta AI, 2026), show that this harness can itself be optimized or even self-rewritten by an agentic proposer -- at the cost of either an expensive code-search loop or unconstrained self-modifying code, neither of which is auditable or usable with a fully black-box model API. We take a narrower, more constrained position: treat the harness as a small, fixed, human-legible action space and learn a policy over it online with classic sample-efficient reinforcement learning (an $ε$-greedy contextual bandit and REINFORCE), scored against a multi-objective reward (task success, verifier score, policy compliance, cost, latency, and an unsupported-claim penalty). We instantiate this control system with DSPy (Khattab et al., 2024) as both the context assembler and the source of the strongest non-adaptive baseline (a DSPy BootstrapFewShot static prompt), and evaluate it across three verifiable task domains -- tool-use workflows, code generation (HumanEval), and multi-hop retrieval QA (HotpotQA) -- and two model providers (a local Ollama model and AWS Bedrock). We release the harness-control-system code, the cross-domain verifiable task suite, the full trajectory/reward-decomposition logs from training, and a provider-agnostic deployment recipe for applying this to a new organization's domain and verification setup.

Comments8 pages, 1 figure, 3 tables. Code and dataset: https://github.com/dpaul0501/context-optimization-rl

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑