arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21515cs.CRcs.LG

ServeGuard:在不泄露被认证读取因子的情况下,对操作者不可见通道的可验证、有界残差限制

ServeGuard: Verifiable, Bounded-Residual Confinement of Operator-Invisible Channels Without Revealing the Certified Read Factor

Dominik Dahlem, Rui Vieira

首次发表
浏览论文内容

中文总结 AI 辅助

ServeGuard通过零知识证明使适配器仅通过监控器覆盖方向读取输入,结构性消除隐藏通道,无需泄露读取因子,实现可验证的供应链安全。

中文摘要 AI 辅助

面向开放权重语言模型的第三方适配器以不透明的权重矩阵形式发布;接收方无法在不信任发布者或检查权重(发布者的核心资产)的情况下检查适配器是否隐藏后门。对于一类重要情况(载荷位于安全监控器结构性盲区内),检测作为防御手段是不可靠的:每个通过声明的监控器进行分解的检测器在其盲子空间上是不变的,并且诚实的适配器与带后门的适配器在我们评估的每个盲子空间统计量上重叠,因为良性适配也使用该子空间。我们不检测这个通道,而是使其在结构上“不存在”并证明我们做到了。发布者构建适配器,使其仅通过监控器覆盖的方向读取输入,并以零知识方式证明这一点,不透露其认证的读取因子。该证书成本低廉,因为昂贵的部分(识别监控器的盲点)是公开基础模型的确定性函数,因此只需证明一个线性恒等式;所服务的残差是基础模型自身的公开下限,而非证明者选择的容差。结果是ServeGuard,一种供应链原语:发布者发布一个“携带证明的适配器”,其证明允许消费者或监管者在不了解被认证读取因子且不信任发布者的情况下,验证该适配器相对于声明的监控器不携带此类隐藏通道;在准入时的类型守卫将保证绑定到服务时准入的适配器字节。在来自四个系列的八个检查点(最高7B)上,监控预算由架构决定:在分组查询检查点上,测量到的前沿在值路径秩处饱和,但在多头检查点上则不然。在0.5B模型上,对良性适配的限制几乎免费,使监控器质量成为安全杠杆。

英文摘要

Third-party adapters for open-weight language models ship as opaque weight matrices; a recipient cannot check whether an adapter hides a backdoor without trusting the publisher or inspecting the weights, the publisher's core asset. For one important class (payloads placed where a safety monitor is structurally blind), detection is unsound as a defense: every detector that factors through the declared monitor is invariant on its blind subspace, and honest and backdoored adapters overlap on every blind-subspace statistic we evaluate, because benign adaptation uses that subspace too. Rather than detect this channel, we make it structurally \emph{absent} and prove that we did. The publisher builds the adapter to read the input only through directions the monitor covers and proves this in zero knowledge, revealing nothing about the read factor it certifies. The certificate is cheap because the expensive part, identifying the monitor's blind spot, is a deterministic function of the \emph{public} base model, so only one linear identity is proved; the served residual is the base model's own public floor, not a prover-chosen tolerance. The result is \emph{ServeGuard}, a supply-chain primitive: the publisher ships a \emph{proof-carrying adapter} whose proof lets a consumer or regulator verify, without the certified read factor and without trusting the publisher, that the adapter carries no hidden channel of this class relative to the declared monitor; an admission-time typing guard binds the guarantee to the adapter bytes admitted at serving time. Across eight checkpoints up to 7B from four families, the monitoring budget is architectural: the measured frontier saturates at the value-path rank on grouped-query checkpoints but not on multi-head ones. On a 0.5B model confinement is nearly free for benign adaptation, making monitor quality the security lever.

发表机构

  • Red Hat AI(红帽人工智能)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑