arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估堆栈而非层:用于智能体动作的确定性与LLM门控是否独立失效?

Evaluate the Stack, Not the Layer: Do Deterministic and LLM Gates for Agent Actions Fail Independently?

Chenglin Yang

arXiv 2610.07359首次发表:更新:

发表机构

University of Lancashire(兰开夏大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过1,119个带标签的智能体动作测试运行时门控堆栈的独立失效假设,发现LLM评判器间耦合显著而规则层与评判器组合近似独立,且服务模型版本不一致可推翻分析结论。

AI 中文摘要

针对智能体工具调用的运行时门控被堆叠起来,其假设是它们的错误会相乘。我们在来自三个语料库的1,119个带标签的智能体动作上对此进行了测试,且没有自适应对手。该堆栈包含一个确定性规则层和四个LLM评判器,其中三个在每次调用时都记录了所服务的模型并重新收集。我们将每个堆栈视为若干乘法等效层(n_mult),其下限在完全耦合下成立。在STRICT漏报定义下(升级给人类记为未阻止),任意两个评判器组合约为1.2至1.4层(φ中位数+0.430,6/6对显著,下限1.02至1.17)。规则层加一个评判器组合为1.86至2.09层(φ中位数+0.014,0/4显著,下限1.01至1.09)。在PRIMARY定义下(升级记为已捕获),区间分别为1.21至1.57和1.80至2.13。在合并数据上区间分离,在每个语料库上点估计分裂,第三方评判器落入评判器区间。单独准确率不能预测某层能增加什么:云规则包将规则层的单独漏报率降低了20%,但未增加新的联合覆盖。评判器耦合的困难份额不可识别:根据探针和漏报定义,在31.8%至61.8%之间。在112个批次中,有50个批次的一个评判器层级由未请求的模型版本服务,集中在外部语料库。该事件推翻了一个预先声明的分析规则,审查评分的逆转导致五个结论反转。我们报告两者。

英文摘要

Runtime gates for agent tool calls are stacked on the assumption that their errors multiply. We test it on 1,119 labelled agent actions from three corpora, without an adaptive adversary. The stack has one deterministic rule layer and four LLM judges, three of them re-collected with the served model recorded on every call. We read each stack as a number of multiplication-equivalent layers, n_mult, with its floor under perfect coupling. Under the STRICT miss definition (escalation to a human scored as not stopped), any two judges compose to about 1.2 to 1.4 layers (ϕ median +0.430, 6 of 6 pairs significant, floors 1.02 to 1.17). The rule layer plus one judge composes to 1.86 to 2.09 layers (ϕ median +0.014, 0 of 4 significant, floors 1.01 to 1.09). Under PRIMARY (escalation scored as caught) the bands are 1.21 to 1.57 and 1.80 to 2.13. Intervals separate on the pooled data, point estimates split on each corpus, and a third-vendor judge lands in the judge band. Solo accuracy does not predict what a layer adds: a cloud rule pack lowers the rule layer's solo miss rate by 20% and adds no new joint coverage. The difficulty share of judge coupling is not identifiable: 31.8% to 61.8% depending on the probe and the miss definition. One judge tier was served by an unrequested model version in 50 of 112 batches, concentrated on the external corpus. That event overturned a pre-declared analysis rule, and the scoring of review verdicts reversed five conclusions. We report both.

Comments15 pages, 1 figure, 10 tables. Artifact (data, scripts, provenance): https://github.com/chenglin1112/evaluate-the-stack

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑