arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06690cs.CRcs.AIcs.LG

策略掩码私有专家:稀疏混合专家(MoE)模型中的可审计可逆能力访问控制

Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models

Zhuoheng Huang, Mukesh Singh

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出Policy-Masked Private Experts方法,在稀疏MoE模型中实现可审计可逆的能力访问控制,经Qwen、DeepSeek等实验验证其能严格限制未授权访问并保留私有分支效用。

中文摘要 AI 辅助

大多数语言模型的访问控制在规范行为的同时,让所有请求都能使用相同的计算资源。我们研究一个不同的系统问题:可信授权能否决定前向传播可访问哪些新训练的参数?策略掩码私有专家(Policy-Masked Private Experts)冻结预训练的稀疏混合专家(MoE)模型,训练一个不相交的专家分支,并在top-k路由前选择公共或私有参数池。其声明的范围狭窄但可验证:在声明的可信计算基(TCB)下,未授权请求不会执行任何私有专家。这并不意味着公共模型缺乏相同的语义能力。我们在Qwen3-30B-A3B和DeepSeek-V2-Lite中测试这种执行控制与任务效用的分离。3个Qwen BF16种子更新所有32个私有专家,同时公共指纹保持不变。在64个对抗场景和96个拒绝/故障关闭事件中,未授权私有执行次数为0;独立钩子精确匹配11616条路由的私有行,且允许-拒绝-允许的恢复完全准确。在两个预先冻结的Qwen基准上,私有分支将精确工具使用的准确率提升5.0个百分点(pp)(分歧数为5对0;单侧Holm检验p=0.03125,对应双侧精确p=0.0625)和21.3个百分点(百分引导法95%置信区间[13.3,29.3],Holm检验p=0.000031)。3名盲态模型评估者测得的外部正效应为18.7个百分点(95%置信区间[9.3,28.0])。参数匹配的LoRA具有相似的外部效用,但事后请求门会使1225个适配器调用被拒绝;而不相交专家分支不会出现此类情况。DeepSeek复现了路由不变性并获得27.0个百分点的提升。有效的密封评估结果接近中性。这些结果支持对训练参数路径进行可审计、可逆的控制,同时表明有用的迁移仍依赖于分布。

英文摘要

Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems question: can trusted authorization determine which newly trained parameters are reachable by the forward pass? Policy-Masked Private Experts freezes a pretrained sparse Mixture-of-Experts (MoE) model, trains a disjoint expert branch, and selects the public or private pool before top-k routing. The resulting claim is narrow but testable: under the declared trusted computing base (TCB), an unauthorized request executes no private expert. It does not imply that the public model lacks the same semantic capability. We test this separation between execution control and task utility in Qwen3-30B-A3B and DeepSeek-V2-Lite. Three Qwen BF16 seeds update all 32 private experts while the public fingerprint remains unchanged. Across 64 adversarial scenarios and 96 deny/fail-closed events, unauthorized private execution is zero; independent hooks exactly match 11,616 routed private rows and allow-deny-allow recovery is exact. On two prospectively frozen Qwen benchmarks, the private branch improves exact tool use by 5.0 percentage points (pp) (five versus zero discordances; one-sided Holm p = 0.03125, corresponding two-sided exact p = 0.0625) and 21.3 pp (percentile-bootstrap 95% CI [13.3, 29.3], Holm p = 0.000031). Three arm-blinded model evaluators retain a positive external effect of 18.7 pp (95% CI [9.3, 28.0]). A parameter-matched Lora has similar external utility, but a post-hoc request gate leaves 1,225 adapter calls under deny; the disjoint expert branch leaves none. DeepSeek reproduces the route invariant and gains 27.0 pp. A valid sealed evaluation is near-neutral. These results support auditable, reversible control over a trained parameter path, while showing that useful transfer remains distribution dependent.

补充信息

↑