AI 中文总结
本研究提出用能力层架建模智能体 harness 故障,通过受控实验验证其可减少候选预算,在SWE-bench多语言库测试中,弃权门解决127个问题,支持其受控不变性机制。
AI 中文摘要
智能体 harness 结合了检索、路由、状态、来源和验证功能,但局部成功的组件可能在共享状态上存在不一致。我们用有限的\textit{能力层架(capability sheaf)}对这种故障进行建模:茎(stalk)编码类型化行为签名,限制映射(restriction map)保留共享字段,而可接受的运行则是有用的全局截面。一个精确的有限约束满足问题(CSP)定义了可接受性,而线性化的相对上同调类则提供了诊断和搜索功能。我们在20个任务集群上进行了受控实验,引入了隐藏的内部中介,其原始状态为干扰变量。对它们的上边缘(coboundary)进行商处理,将每个集群的候选预算从2000个减少到1000个;对齐隐藏状态消除了差距。精确CSP与商匹配,因此该结果证明了对过时代表的不变性,而非优于精确推理。随后,我们在PatchFuseBench的SWE-bench多语言库的发现拆分上测试该方法:来自20个仓库的160个问题、875个真实候选补丁、2579个源感知编辑原子以及153个新执行的补丁。第一个池级构造是恒定的,因为在$\text{coker}D$中$[b-Dx]=[b]$,因此无法对配置进行排名。候选索引修复在875个候选中的848个上是非平凡的,在160个问题中的120个内变化。它解决了118个问题,而匹配的非上同调选择器解决了116个,但该差异在仓库间不显著(精确符号翻转$p=0.75$)。留一仓库的弃权(不执行)门达到127/160,与强锚点持平,比其匹配门多解决1个问题($p=1.0$)。因此,发现门失败,验证拆分仍未突破。该研究支持受控不变性机制和可识别性校正,但不支持真实世界的上同调优势。
英文摘要
Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared state. We model this failure with a finite \emph{capability sheaf}: stalks encode typed behavior signatures, restriction maps retain shared fields, and accepted runs are useful global sections. An exact finite constraint-satisfaction problem (CSP) defines acceptance, while a linearized relative cohomology class provides a diagnostic and search feature. A controlled experiment over 20 task clusters introduces hidden interior mediators whose raw states are nuisance variables. Quotienting their coboundaries reduces the candidate budget from 2,000 to 1,000 per cluster; aligning the hidden state removes the gap. Exact CSP matches the quotient, so the result demonstrates invariance to stale representatives, not superiority over exact reasoning. We then test the method on a discovery split from the SWE-bench Multilingual pool of PatchFuseBench: 160 issues from 20 repositories, 875 real candidate patches, 2,579 source-aware edit atoms, and 153 newly executed patches. A first pool-level construction is constant because $[b-Dx]=[b]$ in $\operatorname{coker}D$ and therefore cannot rank configurations. A candidate-indexed repair is nontrivial on 848/875 candidates and varies within 120/160 issues. It resolves 118 issues versus 116 for a matched noncohomological selector, but the difference is not supported across repositories (exact sign-flip $p=0.75$). A leave-one-repository-out abstention gate reaches 127/160, tying the strong anchor and exceeding its matched gate by one issue ($p=1.0$). The discovery gate therefore fails and the confirmatory split remains sealed. The study supports the controlled invariance mechanism and an identifiability correction, but not a real-world cohomological advantage.