arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Granite.Trust 策略工具:面向生成式 AI 应用的可共享、可执行策略

Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications

Nathalie Baracaldo, Nicolas Mello, Kush R. Varshney, Heiko Ludwig, Kate Soule, David Cox

arXiv 2608.23870首次发表:更新:

发表机构

IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有策略无法适配 GenAI 内容约束执行的问题,提出 Actionable Policy 模式与合成数据流水线等工具,实现 GenAI 全生命周期的策略规范与执行,相关资源开源。

AI 中文摘要

在生成式 AI 的安全策略方面,一刀切的方案并不适用。每个组织和用例都需要根据应用场景、监管环境、组织价值观和用户角色来缓解不同的风险。然而,现有的策略规范方法是为传统访问控制设计的,无法捕捉 GenAI 应用的细微差别:基于内容的约束的执行。我们提出两项贡献来解决这一差距:(1)Actionable Policy 模式,一种基于 YAML 的格式,用于指定模型响应可以包含和不能包含的内容。该模式支持基于例外的策略治理,提出例外以跟踪策略违规情况;(2)合成数据生成流水线,可生成与策略对齐的训练数据,用于模型对齐和测试,以及一组帮助定义该模式和执行策略的工具。这些共同使组织能够一次指定策略,并在 GenAI 应用生命周期中执行:从模型对齐到运行时监控。Actionable Policy 模式、示例策略和工具作为开源提供:[此 URL]。我们欢迎新的想法、贡献和反馈。

英文摘要

When it comes to safety policies for generative AI, one size does not fit all. Each organization and use case needs to mitigate different risks depending on the application context, regulatory environment, organizational values, and user personas. Yet, existing policy specification approaches are designed for traditional access control and fail to capture the nuances of GenAI application: the enforcement of content-based constraints. We present two contributions to address this gap: (1) the Actionable Policy schema, a YAML-based format for specifying what model responses can and cannot contain. The schema enables exception-based policy governance, proposing exceptions to track policy violations; (2) synthetic data generation pipeline that produces policy-aligned training data for model alignment and testing, and a set of tools to help define the schema and enforce policy. Together, these enable organizations to specify policies once and enforce them throughout the GenAI application lifecycle: from model alignment to runtime monitoring. The Actionable Policy schema, example policies, and tools are available as open source: https://github.com/ibm-granite/granite.trust.policy-tools We welcome new ideas, contributions and feedback.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑