发表机构
Tel Aviv University; IITU(特拉维夫大学; 国际信息技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对软件工程 LLM 智能体的错误行动代价高的问题,提出推测不确定性(SU)方法,通过草稿模型生成执行前否决门,降低了 Qwen3-Coder-480B 和 Claude 3.5 Sonnet 的执行错误率与令牌成本,且具备迁移性与通用性。
AI 中文摘要
部署在软件工程领域的大语言模型(LLM)智能体常会产生代价高昂的失败:它们会自信地做出错误行动,而错误行动仅在代价高昂的执行与重试后才被识别。我们提出推测不确定性(Speculative Uncertainty, SU),一种仅从黑盒智能体的输出标记中恢复预测失败信号的方法,无需访问 logits、权重、激活函数或重复采样。通过反转推测解码,一个小型开放权重草稿模型在单次前向传播中对智能体已生成的轨迹进行评分。我们从这些推测交叉似然中提取感知阶段特征,方法是将推理与动作区间分离,并针对可验证目标对其进行校准。SU 生成的失败似然分数可直接被任何下游策略(如路由、人工干预或额外测试时计算)使用。为证明该信号可落地,我们在软件工程智能体 Qwen3-Coder-480B 和闭源 Claude 3.5 Sonnet 上实例化了一种此类策略——执行前否决门,在部署中将执行错误率降低 6-8 个百分点,令牌成本降低 14-19%,无需重新训练即可迁移到分布外基准,并在不同智能体模型间具备通用性。
英文摘要
LLM agents deployed for software engineering fail expensively: they act confidently wrong, and bad actions are recognized only after costly execution and retry. We present Speculative Uncertainty (SU), a method that recovers a predictive failure signal for a black-box agent from its output tokens alone, with no access to logits, weights, activations, or repeated sampling. Inverting speculative decoding, a small open-weight draft model scores the agent's already-generated trajectory in a single forward pass. From these speculative cross-likelihoods we extract phase-aware features by separating the reasoning and action spans, and calibrate them against a verifiable objective. SU produces a failure-likelihood score that any downstream policy, such as routing, human intervention, or extra test-time compute, can consume directly. To show the signal is actionable, we instantiate one such policy, a pre-execution veto gate, on software engineering agents Qwen3-Coder-480B and closed-source Claude 3.5 Sonnet, cutting execution error rate by 6-8 percentage points and token cost by 14-19% in deployment, transferring to out-of-distribution benchmarks without retraining, and generalizing across agent models.
Comments11 pages, 2 figures, EMNLP 2026, Industry track