能力门控语言模型:安全属性可组合,效用则不可
Capability-Gated Language Models: Security Composes, Utility Does Not
浏览论文内容
中文总结 AI 辅助
该研究提出能力门控语言模型,在单一权重内实现按主体的访问控制,证明其安全属性可组合但效用不可组合,为语言模型的安全部署提供了新方案。
中文摘要 AI 辅助
部署的语言模型安全防护措施(安全微调、过滤、遗忘)仅在模型权重之外随主体变化:过滤器需重新配置、层级不断增加、人工制品需重新发布;而在单一权重内部,每个请求都使用相同的模型配置。这促使我们定义能力门控部署:在单一权重内部实现按主体的访问控制,其配置形成格结构—— meet 操作累加主体的限制,join 操作汇集联盟的权限范围。我们通过在现有嵌套分解机制上应用稀疏秩门控实现该部署,用单次归因引导配置文件搜索,并从预注册的保留拆分中读取每个结果。安全属性可组合:在单调 elicitation 假设下,meet 操作可证明能逐点实现误报。在两个谱系中,中位数保留 meet 操作会增强抑制效果,经校正后唯一留存的效应进一步强化了这种抑制。但效用不可组合:单独无害的配置文件组合后会导致保留和流畅性受损,且不存在可组合的效用边界。
英文摘要
Deployed language model safeguards (safety fine-tuning, filtering, unlearning) vary by principal only outside the model weights: filters are reconfigured, tiers are multiplied, and artefacts are reissued; inside one set of weights every request meets the same model configuration. This motivates us to define capability-gated deployment: per-principal access control inside one set of weights, whose configurations form a lattice - meets accumulate a principal's restrictions and joins pool a coalition's reach. We instantiate it by sparse rank gating over an existing nested-factorisation mechanism, guide profile search with one-pass attribution, and read every result once from a pre-registered held-out split. Security approximately composes: provably exactly at meets under a monotone-elicitation assumption we falsify pointwise. In two lineages the median held-out meet deepens suppression; the one effect surviving correction strengthens it. Utility does not: individually harmless profiles can compose to retention and fluency damage, and no compositional bound exists.
发表机构
- BPTI
- askEarth AG(askEarth AG公司)
机构由 AI 辅助整理,请以论文原文为准。