发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出利用低成本开源替代模型的对数概率审计黑盒LLM代理,通过多种读出方式检测错误调用,无需训练即可提升编码任务准确率与任务成功率。
AI 中文摘要
已部署的LLM代理会发出工具调用、查询和代码,这些内容可能悄然出错——当错误浮出水面时,操作已经执行。前沿聊天API隐藏了模型的令牌概率;代理声明的置信度在关键错误上几乎不比随机猜测好;而重采样也无济于事,因为前沿模型高度重复,会在样本间复现相同的调用。我们通过并行运行一个低成本的开源替代模型来恢复缺失的信号。该替代模型读取与代理相同的上下文、模式和提议操作,然后通过一系列互补的读出方式从其自身的对数概率对调用进行评分:教师强制和请求PMI衡量每个参数值的可能性,判别性裁决整体判断调用,工具选择竞争则将函数与其同类进行比较。一个原则决定了信任哪个:生成式似然定位错误的参数值,而裁决捕获整体错误的调用。当错误类型未知时,集成是低遗憾的默认选择。该读出方式无需训练,无需访问代理内部,且每次工具调用仅需一次预填充传递。在困难的编码任务上,它达到了AUROC 0.825,而行动者的声明置信度接近随机(0.598),生成式读出在另外三个行动者上比其高出+0.07至+0.28。与自一致性相比,在近确定性行动者上,它获得了+0.14至+0.19的提升,且成本仅为后者的1/K。该信号驱动两种部署模式:实时门控,将最不可信的调用升级以供审查(在50%覆盖率下,接受操作准确率提升+0.05至+0.30);以及置信度反馈,将工具结果与评分一起返回,以便代理调整其下一步——在实时执行基准上提升任务成功率(+0.119和+0.137,p≤1e-4),并在步骤错误静默时击败随机值控制(+0.078,p=0.003)。
英文摘要
A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier models are highly repetitive, reproducing the same call across samples. We recover the missing signal from a low-cost open-weight surrogate run in parallel. It reads the same context, schema, and proposed action as the agent, then scores the call from its own log-probabilities through a family of complementary readouts: teacher forcing and request-PMI weigh the likelihood of each argument value, a discriminative verdict judges the call as a whole, and tool-choice competition tests the function against its siblings. One principle says which to trust: a generative likelihood localizes wrong argument values, while the verdict catches holistically wrong calls. When the error type is unknown, an ensemble is the low-regret default. The readout is training-free, needs no access to the agent's internals, and costs one prefill pass alongside the tool call. On difficult coding tasks it reaches AUROC 0.825 where the actor's stated confidence is near chance (0.598), and the generative readouts beat it by +0.07 to +0.28 across three further actors. Against self-consistency it gains +0.14 to +0.19 on near-deterministic actors, at 1/K the cost. The signal drives two deployment modes: a real-time gate escalating the least-trustworthy calls for review (+0.05 to +0.30 accepted-action accuracy at 50% coverage), and confidence feedback, returning the tool result with the score so the agent adapts its next step -- lifting task success on live-execution benchmarks (+0.119 and +0.137, p <= 1e-4) and beating a random-value control where step errors are silent (+0.078, p = 0.003).
Comments20 pages, 6 figures, 12 tables