AI 中文总结
该研究针对LLM智能体提出贝叶斯自我升级策略,将生成中委派建模为贝叶斯最优停止问题,经模拟和真实代码级联验证,其升级机制在同等成本下优于事后路由,能力信念区分度随生成提升。
AI 中文摘要
当前的大语言模型(LLM)智能体系统在推理开始前决定是否委派任务(由路由器选择模型),或在响应完成后决定(由验证器对其评分并可能重试)。我们研究第三种机制:智能体在自身推理过程中识别到自身不太可能成功时,将控制权转移给更强的模型。我们将生成中委派问题建模为基于学习到的能力后验的贝叶斯最优停止问题,该后验是对智能体最终任务成功情况的在线估计,其充分统计量从带标签轨迹中学习,而非从原始熵中读取。我们以闭式形式推导了近视升级阈值,通过动态规划刻画了最优策略,并证明该最优策略是随时间变化的阈值,对原始信号无形状假设。我们进一步证明,先验信念在信号的切尔诺夫信息率上呈指数分离,遗憾边界由后验的校准程度决定,以及有限样本保证:使用n个带标签校准轨迹时,部署的插件策略的遗憾以1/√n的速率衰减。受控模拟研究证实了理论的每项预测,包括预测的1/√n速率。我们还报告了在Qwen2.5-Coder 1.5B→7B代码级联(MBPP数据集,257个任务)上的真实模型验证,证实了三个预注册预测中的两个:在同等成本下,升级前沿优于事后路由,且累积能力信念的区分度在生成过程中会提升。
英文摘要
Current LLM agent systems decide delegation before reasoning begins (a router picks a model) or after a response is complete (a verifier scores it and may retry). We study a third regime: an agent that recognises, during its own reasoning, that it is unlikely to succeed and transfers control to a stronger model. We formulate intra-generation delegation as a Bayesian optimal-stopping problem over a learned competence posterior -- an online estimate of the agent's eventual task success whose sufficient statistics are learned from labelled trajectories, not read off raw entropy. We derive the myopic escalation threshold in closed form, characterise the optimal policy via dynamic programming, and prove that the optimal policy is a time-varying threshold with no shape assumption on the raw signal. We further prove exponential separation of the oracle belief at the Chernoff-information rate of the signal, a regret bound governed by the calibration of the posterior, and a finite-sample guarantee: with n labelled calibration trajectories the deployed plug-in policy's regret decays as 1/sqrt(n). A controlled simulation study confirms each prediction of the theory, including the predicted 1/sqrt(n) rate. We additionally report a real-model validation on a Qwen2.5-Coder 1.5B->7B code cascade (MBPP, 257 tasks), confirming two of three pre-registered predictions: the escalation frontier dominates post-hoc routing at equal cost, and the cumulative competence belief's discrimination rises over generation.
Comments21 pages, 5 figures. Code and data: https://github.com/nadeem-shaikh/llm-self-escalation . Associated record: https://doi.org/10.5281/zenodo.21330787