arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FAVA:基于证据支持的权限图的经验证智能体的形式化授权

FAVA: Formal Authorization for Verified Agents with Evidence-Backed Permission Graphs

Yifan Zhang, Xinkui Zhao, Sai Liu, Hengxuan Lou, Guanjie Cheng, Chang Liu

arXiv 2607.27267首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FAVA是一种用于智能体执行的带权限授权框架,通过LLM引导的权限IR、证据支持的权限图和SMT授权器保障安全,在多场景评估中达90.5%决策合规率,可拦截动态违规迹。

AI 中文摘要

大型语言模型(LLM)智能体自主将语义推理与复杂系统操作交织进行。在这些动态环境中,静态工具级权限根本不足;安全授权高度依赖上下文,且严重依赖不断演变的运行时状态和数据流。我们提出FAVA(Formal Authorization for Verified Agents,经验证智能体的形式化授权),这是一种用于智能体执行的带权限的授权框架。FAVA利用LLM引导的权限中间表示(IR)将模糊的自然语言任务转换为结构化约束。随后,确定性降低过程将此IR转换为证据支持的权限图,该图明确跟踪数据流、依赖关系和上下文标签。为提供严格的安全保证,可满足性模理论(SMT)授权器在任何产生效果的操作执行前,都会根据安全策略对当前图进行数学验证。运行时网关随后执行求解器的结果,要么授权执行,要么用精确反例拦截它。我们在OpenAgentSafety、OctoBench和ActPlane场景中评估FAVA。我们的评估表明,FAVA在聚合数据集上达到90.5%的决策合规率(DCR),在评估的迹条件场景中成功拦截动态违规迹。

英文摘要

Large language model (LLM) agents autonomously interleave semantic reasoning with complex system operations. In these dynamic environments, static tool-level permissions are fundamentally insufficient; safe authorization is highly context-dependent and heavily reliant on evolving runtime states and data flows. We present FAVA (Formal Authorization for Verified Agents), a permission-carrying authorization framework for agent execution. FAVA utilizes an LLM-guided Permission Intermediate Representation (IR) to translate ambiguous natural-language tasks into structured constraints. A deterministic lowering pass then converts this IR into an evidence-backed permission graph that explicitly tracks data flows, dependencies, and contextual labels. To provide strict security guarantees, a Satisfiability Modulo Theories (SMT) authorizer mathematically verifies the current graph against security policies before any effectful action executes. A runtime gateway then enforces the solver's result, either authorizing the execution or intercepting it with a precise counterexample. We evaluate FAVA across OpenAgentSafety, OctoBench, and ActPlane scenarios. Our evaluation demonstrates that FAVA achieves a 90.5% Decision Compliance Rate (DCR) over the aggregate dataset, successfully intercepting dynamic violating traces in the evaluated trace-conditioned scenarios.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑