arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

智能体基础设施验证套件中的判别性夹具覆盖

Discriminating Fixture Coverage in Agent-Infrastructure Verification Suites

Xin Xu, Siru Tao

arXiv 2610.02928首次发表:更新:

发表机构

Carnegie Mellon University; Eragon(卡内基梅隆大学; 埃拉贡)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过变异分析评估智能体验证套件的判别力,发现标准验证可能遗漏关键缺陷,提出基于判别性维度覆盖的夹具补充方法,有效提升套件杀死变异体的能力。

AI 中文摘要

不变量套件和运行时监控器日益成为智能体部署决策的门槛,而针对任何特定套件所提供的证据几乎总是一个单一的观察结果:它通过了被认为正确的实现,并使被认为错误的实现失败。我们衡量这一观察结果的价值。将变异分析应用于一个多会话智能体状态投影层的不变量套件,我们首先发现,这种标准验证认证了一个套件,在该套件中,一个移除事件身份去重的一阶变异体通过了所有检查。然后,我们冻结修复后的十二项检查套件,记录其哈希值,并针对由一位未设计任何夹具的对抗性读者指定的十个变异体运行一次:它杀死了五个。对五个幸存者进行参考实现下的插桩显示,它们以两种不同的方式失败,而非一种。三个从未被激活,因为没有夹具提供任何输入,使得变异代码的行为完全不同。另外两个破坏了套件中任何预言机都无法观察到的内部状态。这两种模式需要不同的修复,且两者都无法从通过/失败报告中看出。将缺失输入视为覆盖问题,我们枚举了输入空间的七个判别性维度,预先注册哪些维度未被覆盖以及它们应解释哪些幸存者,并为每个未覆盖维度添加一个夹具,同时逐字重用现有预言机。所有五个幸存者随后都死亡,每个都死于为其预测维度编写的检查。我们将其报告为同一挑战集上的修复结果,而非第二次留出估计,并提供工件,包括冻结的哈希值、注册的预测、所有变异体和运行日志,以便该区分是可检查的。

英文摘要

Invariant suites and runtime monitors increasingly gate agent deployment decisions, and the evidence offered for any particular suite is almost always a single observation: it passes an implementation believed correct and fails one believed broken. We measure what that observation is worth. Applying mutation analysis to an invariant suite for a multi-session agent state-projection layer, we first find that this standard validation certifies a suite in which a first-order mutant removing event-identity deduplication survives every check. We then freeze the repaired twelve-check suite, record its hash, and run it once against ten mutants specified by an adversarial reader who designed none of its fixtures: it kills five. Instrumenting the five survivors against the reference shows they fail in two distinct ways, not one. Three are never activated, because no fixture supplies an input on which the mutated code behaves differently at all. The other two corrupt internal state that no oracle in the suite can observe. The two modes need different repairs, and neither is visible from a pass/fail report. Treating the missing inputs as a coverage question, we enumerate seven discriminating dimensions of the input space, register in advance which are uncovered and which survivors they should explain, and add one fixture per uncovered dimension while reusing the existing oracles verbatim. All five survivors then die, each to the check written for its predicted dimension. We report this as a repair result on the same challenge set rather than a second held-out estimate, and give the artifact, including the frozen hash, the registered predictions, all mutants and the run logs, so the distinction is checkable.

Comments10 pages, 1 figure, 2 tables. NeurIPS 2026 Workshop: Who Verifies the Agents? Toward Reliable Agent Development

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑