发表机构
Blossom AI; Blossom AI Labs(Blossom人工智能; Blossom人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现验证器可能泄露答案,使优化器比较失效;提出支持门控验证契约,通过参考门和运行时门确保证据资格,预注册实验验证了其有效性。
AI 中文摘要
智能体开发者越来越多地通过模拟器验证器来比较提示、工具、策略和诊断算法。然而,验证器可能使求解器比较变得无意义:如果其探测或谓词编码了目标身份,精确优化器可能看似有效,而实际上并未解决任何真正的歧义。我们在一个闭环决策智能体的聚合轨迹调试器中报告了这种失败。精确最小碰集(MHS)和一种传播感知的贪心方法在12个开发案例中的12个返回了相同的支持集,并在9/12个案例中实现了相同的植入故障恢复。随后的审计发现,精确锚点谓词在9/9个案例中产生了植入对。移除这些锚点后,整体植入对恢复率为8/9;然而,硬探测单例对在9/9个案例中与植入对匹配,且没有案例在传播后保留非空残差冲突族(0/9)。优化器是正确的,但验证器已经泄露了答案。我们用支持门控验证契约取代了求解器优先的评估。首先,干净的参考图必须显示重复的组件暴露;然后,匹配的参考/当前门必须建立可比较的运行时证据;只有在此之后,独立校准的信号规则才能返回检测结果。在一个包含1,440个案例和21,600个分区行的预注册保留集中,55/72个制度-组件单元通过了参考门,54/55个通过了运行时门,稳定误接纳为0/20个代表性组件,单侧精确95%上限为0.1391。在接纳的单元内,受影响的干净流量比名义故障单元比例更好地预测了检测。主要教训是结构性的:在优化组件选择器之前,验证证据的资格性和非泄露性。否则,更强的求解器只能证明更强的验证器伪影。
英文摘要
Agent developers increasingly compare prompts, tools, policies, and diagnosis algorithms through simulator-grounded verifiers. A verifier can nevertheless make a solver comparison vacuous: if its probes or predicates encode the target identity, an exact optimizer may appear effective without resolving any genuine ambiguity. We report such a failure in an aggregate-trace debugger for a closed-loop decision agent. Exact minimum hitting set (MHS) and a propagation-aware greedy method returned identical supports in 12/12 development cases and the same planted-fault recovery in 9/12. A subsequent audit found that exact-anchor predicates produced the planted pair in 9/9 cases. After removing those anchors, overall planted-pair recovery was 8/9; hard-probe singleton pairs nevertheless matched the planted pair in 9/9, and no case retained a nonempty residual conflict family after propagation (0/9). The optimizer was correct, but the verifier had already disclosed the answer. We replace solver-first evaluation with a support-gated verification contract. A clean reference map must first show repeated component exposure; a matched reference/current gate must then establish comparable runtime evidence; only afterward may an independently calibrated signal rule return a detection. In a preregistered heldout comprising 1,440 cases and 21,600 partition rows, 55/72 regime-component units passed the reference gate, 54/55 passed the runtime gate, and stable false admission was 0/20 represented components with a one-sided exact 95% upper bound of 0.1391. Within admitted units, affected clean traffic predicted detection better than nominal fault-cell fraction. The main lesson is structural: verify evidence eligibility and non-revelation before optimizing the component selector. Otherwise a stronger solver can merely certify a stronger verifier artifact.
CommentsSubmitted to Who Verifies the Agents? Toward Reliable Agent Development (NeurIPS 2026 workshop). 7 pages, 0 figures, 2 tables. The reproducibility artifact is linked in the paper