arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CAPO:面向大语言模型智能体的感知约束提示优化

CAPO: Constraint-Aware Prompt Optimization for LLM Agents

Victor Ye Dong, Reid Pryzant, Yi Liu, Jian Jiao

arXiv 2608.16068首次发表:更新:

发表机构

Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出CAPO与DCAPO两种方法,通过原始对偶框架优化LLM智能体的系统提示,在智能体及助手式任务中均提升了可行性与性能。

AI 中文摘要

大语言模型(LLM)正越来越多地被部署为依赖系统提示来使用工具和完成任务的智能体。这类部署带来了明确的操作要求,包括恰当的工具使用、简洁的提示与解决方案路径,以及遵守安全和格式政策。然而,对许多从业者而言,收集特定领域的监督数据来对模型进行后训练以满足这些要求是不可行的。我们提出了CAPO(Constraint-Aware Prompt Optimization,感知约束提示优化),一种结合了基于池的重写与自适应约束加权的原始对偶方法,用于在显式操作约束下优化系统提示。在多个智能体基准测试中,CAPO能更可靠地达到经验上可行的操作点,同时提升任务性能。CAPO还能泛化到智能体设置之外,在具有输出格式和安全/隐私约束的助手式评估中取得优异结果。我们进一步提出了DCAPO(Dynamically Trained CAPO,动态训练的CAPO),该方法在冻结任务智能体的同时,使用基于池的GRPO训练一个反馈和对偶条件的重写器。在不同规模的任务智能体上,DCAPO在所有评估领域都生成了可行的提示,且达到或超过了所评估基线的任务准确率。一项替代分析刻画了有限池和离散重写误差如何进入不精确的原始对偶过程。

英文摘要

Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployments impose distinct operational requirements, including appropriate tool use, concise prompts and solution paths, and compliance with safety and formatting policies. For many practitioners, however, assembling domain-specific supervised data to post-train models to meet these requirements is infeasible. We introduce CAPO (Constraint-Aware Prompt Optimization), a primal-dual method that combines pool-based rewrites with adaptive constraint weighting to optimize system prompts under explicit operational constraints. Across agentic benchmarks, CAPO more reliably reaches empirically feasible operating points while improving task performance. CAPO also generalizes beyond agentic settings, achieving strong results on assistant-style evaluations with output-format and safety/privacy constraints. We further introduce DCAPO (Dynamically Trained CAPO), which trains a feedback- and dual-conditioned rewriter with pool-based GRPO while keeping the task agent frozen. Across task agents of different sizes, DCAPO produces a feasible prompt in every evaluated domain and matches or improves the task accuracy achieved by the evaluated baselines. A surrogate analysis characterizes how finite-pool and discrete-rewrite errors enter the inexact primal-dual procedure.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑