发表机构
TraslaIA(特拉斯拉AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体在持久停滞时无法决定恢复或停止的治理困境,提出问责证明块(APB)机制,通过密码学签名绑定证据与决策,实现可问责的权力转移,并经实验验证其完备性与安全性。
AI 中文摘要
一个治理得当的LLM智能体可能达到一种状态,在该状态下继续执行和自动停止均不被允许:系统已检测到其可观测性或漂移检测层存在持续性故障,但无法自行决定谁有权恢复、拒绝或重新校准部署。我们将此称为身份绑定治理事件,并形式化描述解决该事件的机制。我们引入问责证明块(APB):一个系统构建的证据块、一个由人工提供的决策块,以及一个将两者绑定到已注册主体的ed25519签名。系统无法伪造签名,主体也无法在未被察觉的情况下篡改证据。我们证明了四个定理:协议约束治理完备性、不可否认性、匿名重新授权的不可能性以及有限时间内APB构造终止性。实现采用RFC 8785 JSON规范化及基于UUID4的重放谓词。实验上,治理完备性在3812次停滞事件中成立,零未解决案例;验证器在9个对抗向量中检测出100%的1800次攻击;k-of-n多主体变体在2000次单密钥捕获尝试中显示零误接受。对六个开源LLM的研究发现,漂移阈值T*在模型内稳定(sigma/T*<2%),但在模型间存在差异,否定了规模单调性:最大的模型并未发生漂移。因此,T*必须针对每次部署进行测量,而APB是使该阈值产生可问责权力转移的载体。
英文摘要
A correctly governed LLM agent can reach a state in which neither continuing execution nor automatically halting is admissible: the system has detected a persistent failure of its observability or drift-detection layer, but cannot itself decide who has the authority to resume, deny, or recalibrate the deployment. We call this an identity-bound governance event and formalise the mechanism that resolves it. We introduce the Accountability Proof Block (APB): a system-constructed evidence block, a human-supplied decision block, and an ed25519 signature binding both to a registered principal. The system cannot forge the signature, and the principal cannot alter the evidence undetected. We prove four theorems: Protocol-Bounded Governance Completeness, Non-Repudiability, Impossibility of Anonymous Re-Authorization, and Finite-Time APB Construction Termination. The implementation uses RFC 8785 JSON canonicalization and a UUID4-based replay predicate. Empirically, governance completeness holds over 3,812 halt events with zero unresolved cases; the verifier detects 100% of 1,800 attacks across 9 adversarial vectors; a k-of-n multi-principal variant shows 0 false acceptances in 2,000 single-key capture attempts. A study of six open LLMs finds the drift threshold T* stable within model (sigma/T* < 2%) but varying across models, refuting size-monotonicity: the largest model did not drift. T* must therefore be measured per deployment, and the APB is the vehicle by which that threshold yields accountable authority transfer.
Comments24 pages, 7 tables, 2 propositions. Paper 8 of the Agent Governance Series. v2: RFC 8785 canonicalization, event_id uniqueness (V5), k-of-n multi-principal threshold governance (Prop 8.5), infrastructure fault injection experiment. Code: https://github.com/chelof100/identity-bound-governance