arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAFESHIELD:面向小型语言模型部署时安全的决策组织框架

SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Models

Xingru Zhou, Luis Sentis, Aarti Choudhary

arXiv 2610.07276首次发表:更新:

发表机构

The University of Texas at Austin; Advanced Micro Devices, Inc. (AMD)(德克萨斯大学奥斯汀分校; 超威半导体公司(AMD))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对部署时安全决策缺乏组织协调的问题,提出SAFESHIELD框架,通过责任分解与显式协调组织安全决策,实验证明协调机制对端到端安全至关重要。

AI 中文摘要

语言模型的部署时安全通常通过运行时护栏实现,例如输入审核、路由、检索验证和输出过滤。现有的部署框架为这些功能提供了日益强大的机制,但对于它们所产生的安全决策应如何明确组织、协调和审计,提供的指导有限。我们将部署时安全表述为一个决策组织问题,包含两个要素:安全决策的责任导向分解以及它们之间的明确协调。我们在SAFESHIELD中实例化了这一表述,这是一个面向小型语言模型的部署时安全系统,它组织了四个反复出现的决策责任(准入、路由、证据和放行),并将已提交的决策记录在可审计的决策轨迹中。我们通过机制级实验、聚合阶段消融、受控协调消融和面向部署的压力测试套件对SAFESHIELD进行了评估。机制级结果表明,实例化的护栏提供了决策过程所需的能力,而聚合消融显示,随着周围安全组织的移除,端到端安全性大幅下降。更重要的是,专门的协调消融在保留参与护栏机制的同时,选择性地切断它们的依赖关系:移除准入门控显著增加了错误放行,而从放行决策中扣留上游证据使放行准确率从96.0%降至69.5%。这些结果提供了系统级证据,表明部署时安全不仅取决于单个护栏的能力,还取决于它们的决策如何被组织和协调。

英文摘要

Deployment-time safety of language models is commonly implemented through runtime guardrails such as input moderation, routing, retrieval verification, and output filtering. Existing deployment frameworks provide increasingly capable mechanisms for these functions, but offer limited guidance on how the safety decisions they produce should be explicitly organized, coordinated, and audited. We formulate deployment-time safety as a decision-organization problem with two elements: responsibility-oriented decomposition of safety decisions and explicit coordination among them. We instantiate this formulation in SAFESHIELD, a deployment-time safety system for small language models that organizes four recurring decision responsibilities (admission, routing, evidence, and release) and records committed decisions in auditable Decision Traces. We evaluate SAFESHIELD through mechanism-level experiments, aggregate stage ablations, controlled coordination ablations, and a deployment-oriented stress suite. Mechanism-level results show that the instantiated safeguards provide the capabilities required by the decision process, while aggregate ablations show substantial degradation in end-to-end safety as the surrounding safety organization is removed. More importantly, dedicated coordination ablations preserve the participating safeguard mechanisms while selectively severing their dependencies: removing admission gating substantially increases false release, and withholding upstream evidence from the release decision reduces release accuracy from 96.0% to 69.5%. These results provide system-level evidence that deployment-time safety depends not only on the capability of individual guardrails, but also on how their decisions are organized and coordinated.

CommentsAccepted to the Application Track of IEEE TPS 2026. 12 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑