智能体行为契约II:不假设独立性的组合可靠性验证
Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
浏览论文内容
中文总结 AI 辅助
该研究针对多智能体系统组合可靠性的独立性假设开展验证,发现组件共享模型时正相关性会高估冗余度,提出无依赖假设的有限样本验证方法并通过扩充矩函数提升验证效果,配套任意时间有效验证方法。
中文摘要 AI 辅助
多智能体系统的组合可靠性界通常通过将各组件可靠性相乘得到,该步骤基于条件独立性假设,虽常被提及却极少被验证。我们对此展开验证:在一项预注册评估中,基于确定性代码对18000个任务评分、未使用大语言模型(LLM)评判,某一模型的两个实例在两智能体交接任务中,任一实例失败的任务里,90.0%的任务会共同失败(对数或为6.66,95%置信区间[6.38,7.00];φ值为0.916)。替换为不同模型时,6次对比中有6次关联度降低;替换为不同供应商(其模型已存在差异)时则无此效果,该预注册假设被报告为零结果。此误差有明确指向,且对操作者不利:正相关性会使联合失败率高于独立性乘积,即当组件共享同一模型时,冗余度被过度高估。无独立性假设的替代方案常无实际意义,而拟合依赖模型更差:我们证明,随着样本量n增长,拟合模型函数的自举界对真值的覆盖率会下降,识别差距为O(1),而自举缩减量为O(n^{-1/2});更多数据会使此类验证结果更差,且无明显症状。我们提出一种不假设依赖结构的有限样本验证方法:基于联合分布的线性规划,围绕实测协同执行矩的Bonferroni-Clopper-Pearson区间构建,该方法可靠、对所提供信息尖锐且随矩族单调变化。将10个矩函数扩充至14个,可使识别区间缩小85.7%,并将验证下限从0.2455提升至0.4116。配套的任意时间有效验证方法在可选停止下将I类错误控制在0.0471。常见依赖统计量受边际约束,当对比智能体失败率不同时,可能反转条件的表观排序。本文已发布契约、评分代码、分析脚本及预注册内容。
英文摘要
Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.
发表机构
- Qualixar
机构由 AI 辅助整理,请以论文原文为准。