arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27087cs.AIcs.SC

政策即技能:具有证据、确定性控制和审计的受治理LLM决策支持

Policy-as-Skill: Governed LLM Decision Support with Evidence, Deterministic Control, and Audit

  • Institute for Smart Systems Technologies(智能系统技术研究所)
  • University of Klagenfurt(克拉根福大学)
  • Faculté Polytechnique Université de Kinshasa(金沙萨大学理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Kabeh Mohsenzadegan, Vahid Tavakkoli, Kyandoghere Kyamakya

中文总结 AI 辅助

本文提出Policy-as-Skill(PaS)模块化运行时,将证据验证、审查路由、版本控制和审计打包为可执行政策能力,在600个任务上优于LLM+RAG,并发现确定性控制应选择性应用而非普遍干预。

中文摘要 AI 辅助

组织越来越多地将LLM用于政策、合规、风险和运营决策支持,这需要证据验证、审查路由、版本控制和可审计性。我们引入了政策即技能(Policy-as-Skill,PaS),这是一个模块化运行时,将这些功能打包为可执行、可版本化的政策能力。使用固定的Gemma4后端,在600个开发任务上评估了十三种方法。PaS+审计实现了53.8%的精确准确率、宏F1 0.346、审查F1 0.854、引用精确率1.000、政策参考召回率0.984和审计完整性1.000,在大多数治理和审查指标上优于LLM+RAG。确定性控制将总体准确率提升至61.2%,但强烈依赖任务,支持选择性而非普遍的基于规则的干预。

英文摘要

Organizations increasingly use LLMs for policy, compliance, risk, and operational decision support, requiring evidence validation, review routing, version control, and auditability. We introduce Policy-as-Skill (PaS), a modular runtime that packages these functions as executable, versioned policy capabilities. Thirteen methods are evaluated with a fixed Gemma4 backend on 600 development tasks. PaS+Audit achieves 53.8% exact accuracy, macro-F1 0.346, review F1 0.854, citation precision 1.000, policy-reference recall 0.984, and audit completeness 1.000, outperforming LLM+RAG on most governance and review metrics. Deterministic control raises aggregate accuracy to 61.2% but is strongly task dependent, supporting selective rather than universal rule-based intervention.

↑