发表机构
Stevens Institute of Technology(史蒂文斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RAISE通过两阶段训练(验证监督微调和强化学习)从符号验证信号中学习,使Qwen3.5-9B在Cedar策略合成上超越GPT-6 Astra和Claude Opus 5,显著提升语义成功率。
AI 中文摘要
将自然语言访问控制需求翻译成策略需要仔细推理权限、约束和异常,即使是前沿LLM也常常产生违反预期授权语义的策略。我们构建了CedarInstruct,据我们所知,这是第一个同时支持可形式验证的Cedar策略合成的训练和语义评估的数据集。它包含44个领域的5,800个场景和1,408个代表单一合成组织的场景,每个场景都有一个经过验证的目标策略和一个可执行的验证计划。在此数据上,我们引入了RAISE,它通过两个阶段的正式验证来训练策略合成器:先进行验证监督微调(SFT),然后进行从验证器信号中学习的强化学习(RL)阶段。我们发现,SFT在很大程度上成功是因为让模型表达了它们已有的授权逻辑,因为未经训练的模型很少写出有效的Cedar,但一旦写出,它们往往能正确推理。在SFT之后,验证器信息的使用方式比使用量更重要。在六种消耗越来越丰富的验证器信号的RL实例化中,只有RAISE-OC在SFT上取得了有意义的改进;它将失败的检查和符号反例转化为引导式探索,并通过离上下文GRPO从结果中学习。使用约5.4K个验证场景和LoRA微调,RAISE-OC训练Qwen3.5-9B在保留场景的语义成功率上分别超过零样本GPT-6 Astra和Claude Opus 5达13.33和16.26个百分点,并且训练迁移到独立构建的CedarBench。
英文摘要
Translating natural-language access-control requirements into policies requires careful reasoning about permissions, constraints, and exceptions, and even frontier LLMs often produce policies that violate the intended authorization semantics. We construct CedarInstruct, to our knowledge the first dataset that supports both training and semantic evaluation for formally verifiable Cedar policy synthesis. It contains 5,800 scenarios across 44 domains and 1,408 representing a single synthetic organization, each with a verified target policy and an executable verification plan. On this data we introduce RAISE, which trains policy synthesizers from formal verification in two stages, verified supervised fine-tuning (SFT) followed by a reinforcement learning (RL) stage that learns from verifier signal. We find that SFT succeeds largely by letting models express authorization logic they already have, since untrained models rarely write valid Cedar but often reason correctly when they do. After SFT, how the verifier's information is used matters more than how much of it is used. Of six RL instantiations that consume progressively richer verifier signal, only RAISE-OC improves meaningfully on SFT; it turns failed checks and symbolic counterexamples into guided exploration and learns from the result with off-context GRPO. With about 5.4K verified scenarios and LoRA fine-tuning, RAISE-OC trains Qwen3.5-9B to surpass zero-shot GPT-6 Astra and Claude Opus 5 by 13.33 and 16.26 percentage points in semantic success on held-out scenarios, and training transfers to the independently constructed CedarBench.
Comments34 pages, 5 figures, 11 tables. Code: https://github.com/Aizhouym/raise