arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25408cs.AI

作为受控变量的上下文组装:冻结大语言模型智能体利用策略的控制理论视角

Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents

Debjyoti Paul

首次发表
浏览论文内容

中文总结 AI 辅助

研究聚焦大语言模型智能体,以往控制工具选择等,本文将上下文组装视为受控变量,通过上下文博弈或强化学习策略在线学习,进行形式分解、稳定性论证及不确定性校准分析,给出控制理论视角的相关证据。

中文摘要 AI 辅助

越来越多2026年的工作将控制理论应用于大语言模型智能体,如工具介导控制器的李雅普诺夫认证稳定性、大规模离散工具宇宙上稀疏策略的样本复杂度界限等。本文并不声称将控制理论引入大语言模型智能体,而是关注受控变量是什么。以往工作控制工具选择等,本文将上下文组装本身视为受控变量,由冻结模型外的上下文博弈或强化学习策略在线学习。本文进行了形式分解,给出在线控制器稳定性论证,并报告控制器自身对实际任务结果置信度的不确定性校准分析。应用部分在三个领域和两个模型提供商上实例化相同控制器并发布数据集等,本文聚焦形式框架及控制理论主张所需的稳定性/不确定性证据。

英文摘要

A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al., "Stable Agentic Control", 2026), sample-complexity bounds for sparse policies over massive discrete tool universes (Majumdar, "Sparse Agentic Control", 2026), and regulatory-control decompositions of multi-agent systems into auditable feedback loops (Nogueira and Skogestad, 2026). We do not claim to introduce control theory to LLM agents -- that ship has sailed. Our narrower claim is about what the controlled variable is. Prior work controls tool selection, inter-agent message routing, or the agent's raw action stream. We instead treat context assembly itself -- which prompt template, which few-shot demonstrations, how much retrieved context, how many planning/verification passes -- as the controlled variable, learned online by a contextual bandit or REINFORCE policy sitting outside a frozen model. This paper develops the formal decomposition (inner frozen policy $π_θ$, outer context policy $π_ϕ$), gives a stability argument for the online controller in the sense used by Zhang et al. (2026) (non-decreasing expected reward under bounded policy change), and reports an uncertainty-calibration analysis of the controller's own confidence against realized task outcomes. The applied counterpart to this paper instantiates the same controller across three domains and two model providers and releases the dataset, trajectory logs, and a deployment recipe; here we focus on the formal framing and the stability/uncertainty evidence a control-theoretic claim requires.

补充信息

↑