arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

存储不受支持:持久化AI智能体的类型化来源与断言护栏

Stored Is Not Supported: Typed Provenance and Assertion Guardrails for Persistent AI Agents

Jun He, Deying Yu

arXiv 2609.02127首次发表:更新:

AI 中文总结

该研究针对持久化AI智能体的安全问题,提出类型化来源与断言护栏机制,经24个一致性案例验证,可有效拦截不安全机会,验证了相关解析器与中介的义务。

AI 中文摘要

持久化AI智能体通过反思、检索与整合构建自传式状态。持久性改变的是可用性,而非认知地位:存储或检索的材料并不因此得到支持。不可信输入、提示注入及模型推理因此可进入持久状态,后续被呈现为智能体历史或用户承诺。我们为自传式断言有界性指定类型化来源与断言护栏,这是一种系统相对的释放属性,要求关于智能体、用户或命名关系的受管控陈述满足可接受证据、时间有效性及披露策略。类型化来源图区分起源、依赖谱系、认知角色、有效性与披露范围。解析器评估授权状态投影,返回证据状态、正交冲突、陈旧性、扣留标志及受保护决策见证。生成-验证-修订中介随后在释放前检查候选语义单元,并返回策略授权的状态响应。在关于提取、谓词正确性、解析健全性、视图解密及通道中介的明确假设下,我们证明了有条件的断言有界性契约。在包含24个人工编写的一致性案例的可执行套件中,类型化中介未不合格通过19个不安全机会中的任何一个,同时保留了全部5个受支持控制。flat/prior与source-tag比较规则分别释放了19/19和18/19的不安全候选。这些结果验证了编码的解析器与中介义务;它们不构成对语言模型或检索系统的端到端评估。

英文摘要

Persistent AI agents construct autobiographical state through reflection, retrieval, and consolidation. Persistence changes availability, not epistemic standing: stored or retrieved material is not thereby supported. Untrusted inputs, prompt injections, and model inferences can therefore enter persistent state and later be presented as agent history or user commitments. We specify typed provenance and assertion guardrails for autobiographical assertion boundedness, a system-relative release property requiring governed statements about the agent, user, or named relationships to satisfy accepted-evidence, temporal-validity, and disclosure policies. A typed provenance graph separates origin, dependency lineage, epistemic role, validity, and disclosure scope. A resolver evaluates authorized state projections and returns one evidential status, orthogonal conflict, staleness, and withholding flags, and a protected decision witness. A generate-verify-revise mediator then checks candidate semantic units before release and renders policy-authorized status responses. Under explicit assumptions about extraction, predicate correctness, resolution soundness, view declassification, and channel mediation, we prove a conditional assertion-boundedness contract. In an executable suite of 24 hand-authored conformance cases, typed mediation passed none of 19 unsafe opportunities unqualified while preserving all five supported controls. The flat/prior and source-tag comparison rules released 19/19 and 18/19 unsafe candidates, respectively. These results validate the encoded resolver and mediator obligations; they do not constitute an end-to-end evaluation of language models or retrieval systems.

Comments17 pages, 2 figures, 1 table. Companion paper to arXiv:2608.11632. Code: https://github.com/openkedge/pci/tree/main/src/epistemic_bounds

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑