发表机构
Gaoling School of Artificial Intelligence, Renmin University of China; Institute of Computing Technology, Chinese Academy of Sciences; Duke University; Institute of Automation, Chinese Academy of Sciences; College of Computer Science, Zhejiang University(中国人民大学人工智能学院; 中国科学院计算技术研究所; 杜克大学; 中国科学院自动化研究所; 浙江大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究大语言模型智能体多步任务中早期失败预测问题,提出通过召回控制的探测级联方法,能早期预测失败并转化为实用终止级联,在TextCraft模型上达到高召回目标,节省大量推理计算资源,还刻画了样本复杂度。
AI 中文摘要
解决多步任务的大语言模型(LLM)智能体经常会陷入注定失败的轨迹,并且在失败变得明显之前继续消耗大量推理计算资源。我们表明,从智能体的内部表示中可以早期预测失败:对隐藏激活进行轻量级的每轮探测能够早在第一轮交互中就预测最终情节的失败,而仅读取智能体可观察行为的评分器几乎无法比随机猜测做得更好。我们将这个信号转化为一个实用的终止级联:每轮一个无分布校准门,联合搜索每轮的召回预算,以便最终成功的情节以用户指定的全局速率通过所有门;这种情节级别的保证在部署中很重要,因为错误终止风险会在各个门之间累积。在TextCraft上的两个智能体模型中,级联达到了从90%到97%的每个召回目标,在90%的目标下,节省了Qwen - 2.5 - 7B模型47.1%±10.3%以及Llama - 3.2 - 3B模型37.2%±8.8%的推理计算资源,是最佳单门策略的1.6 - 1.7倍。仅读取行为的相同级联节省的资源约为一半,并且在探测中添加行为特征不会带来进一步的收益:隐藏状态捕捉了行为所揭示的信息。最后,我们刻画了验证高召回目标的样本复杂度,告知从业者他们的数据能够以及无法证明支持哪些召回承诺。代码即将发布。
英文摘要
Large language model (LLM) agents often waste inference compute by continuing multi-step trajectories that are already doomed to fail. We study early failure prediction and inference-time early stopping for LLM agents using hidden-state probes. Lightweight linear probes on internal activations predict eventual task failure from the first interaction round, substantially earlier than agent-monitoring methods based only on observable behavior. We turn this signal into a recall-controlled abort cascade for reducing LLM agent inference costs. The cascade applies a distribution-free calibrated failure detector at each early interaction round and jointly optimizes per-round recall budgets. This design ensures that eventually successful episodes survive all early-stopping gates at a user-specified global recall rate. After selection, the cascade is frozen and certified on independent data, providing an exact post-selection recall guarantee. We evaluate the method on TextCraft and WebShop with Qwen-2.5-7B, Llama-3.2-3B, and Qwen3-1.7B. The proposed LLM agent early-stopping cascade outperforms the best single-gate baseline in every model-environment pair, saving 1.5-8.8 times more compute at a 90% recall target. Achieved recall remains within one standard deviation of its target in all 24 configurations. The strongest settings reduce generated tokens by 60.2% on TextCraft and 54.9% on WebShop at 90% recall, while retaining savings of 45.0% and 41.5% at 95% recall. Behavior-only monitoring is consistently weaker, and adding behavioral features to hidden-state probes provides no further gain. We also characterize the sample complexity required to certify high-recall early-stopping policies. The code will be released soon.