AI 中文总结
PRIMUS提出多智能体联邦治理框架,结合素数幂身份与BLS签名,将误报终止率降至0%,并探索验证器作为分级适应度信号,校准良好但泛化受限。
AI 中文摘要
多智能体联邦需要治理机制,以在对抗条件下回答三个问题:谁参与了(身份)、他们是否合规(执行)以及谁来决定(权威)。另一个问题是,监督联邦输出的验证机制能否同时引导生成-测试循环走向更好的答案。第一部分:PRIMA引入了素数幂智能体身份和一种共识令牌,其因子分解可索引参与情况,但假设智能体是诚实的。我们提出PRIMUS,它将素数幂身份与BLS聚合签名(PIAC)相结合,推导出安全终止阈值,在10%信道噪声下将误报智能体终止率从80%降至0.00%,给出了封闭形式的经济边界,其中单例治理优于拜占庭法定人数($\gamma^* \approx 9f$,在n=50至10,000范围内验证为平坦),并指定了带租约和围栏的VRF继任机制,使安全性在部分同步下无条件成立。五个问题被确定为在该模型内无法修复,并作为范围边界陈述。第二部分:验证器不是求解器。我们询问PRIMA的二元工件保真度判定能否转换为分级适应度信号,并在二元覆盖码上测量转换效果。针对注入故障负担的校准很强(确定性$\rho$=0.676,完整$\rho$=0.819);针对真实LLM生成的候选,相同分数降至0.158和0.406,约为校准值的四分之一(测得的同设计者混淆)。作为预过滤器,它明显优于随机分数对照,并略优于二元门控。在400次显式优化迭代下,它未被博弈,但仅因为目标在一个诚实答案后饱和。跨族评判器保留了负担排序信号,同时破坏了单个判断。未产生覆盖码记录。实测程序成本:164.78美元。
英文摘要
Multi-agent federations need governance that answers three questions under adversarial conditions: who participated (identity), did they conform (enforcement), and who decides (authority). A separate question is whether the verification machinery that polices a federation's outputs can also steer a generate-and-test loop toward better answers. Part I. PRIMA introduced prime-power agent identity and a consensus token whose factorization indexes participation, but assumed honest agents. We present PRIMUS, which couples prime-power identity with BLS aggregate signatures (PIAC), derives a safe-kill threshold that reduces false-positive agent termination from 80% to 0.00% under 10% channel noise, gives the closed-form economic boundary where singleton governance outperforms Byzantine quorum ($γ^* \approx 9f$, verified flat across n = 50 to 10,000), and specifies VRF succession with lease and fencing that makes safety unconditional under partial synchrony. Five problems are identified as provably unfixable within the model and stated as scope boundaries. Part II. A verifier is not a solver. We ask whether PRIMA's binary artifact-fidelity verdict can be converted into a graded fitness signal, and measure the conversion on binary covering codes. Calibration against injected fault burden is strong ($ρ$ = 0.676 deterministic, 0.819 full); against real LLM-generated candidates the same scores fall to 0.158 and 0.406, roughly a quarter of the calibration value (the same-designer confound, measured). As a pre-filter it beats a random-score control convincingly and a binary gate narrowly. Under 400 iterations of explicit optimization it was not gamed, but only because the objective saturated after one honest answer. A cross-family judge preserves the burden-ordering signal while destroying individual judgments. No covering-code record resulted. Measured program cost: USD 164.78.
Comments17 pages, 6 tables, no figures. Extends PRIMA (arXiv:2605.24775). Per-step result files, preregistrations, and manifests available to reviewers on request; certain implementation constants held under controlled release (see Appendix D)