arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16403cs.CRcs.LG

实现随机傅里叶特征的白盒不可检测后门

Implementing a White-Box Undetectable Backdoor for Random Fourier Features

Michael Collins, Jada Cumberland, Brianne Dunn, Ross Gore, Samuel Jackson, Sachin Shetty

首次发表
浏览论文内容

中文总结 AI 辅助

本文用numpy和scipy实现了Goldwasser等人提出的白盒CLWE-RFF后门构造,验证其可用普通科学计算工具实现,并通过统计测试证明后门模型与干净模型不可区分。

中文摘要 AI 辅助

Goldwasser等人表明,在随机傅里叶特征(RFF)算法训练的机器学习模型中,可以在与连续学习误差(CLWE)问题相关的硬度假设下植入不可检测的后门。在标准密码学假设下,即使对模型权重进行完全的白盒审计也无法检测到此类后门。该构造以密码学归约和概率引理的形式陈述,没有参考实现,并依赖于次要机制,如稀疏高斯煎饼分布和齐次CLWE条件密度。仅从论文本身来看,其在普通数值代码中的可实现性并不明显。本文仅使用numpy和scipy端到端实现了白盒CLWE-RFF后门构造,以测试这种威胁是否可通过商品科学计算工具实现,还是需要专门的密码学基础设施。我们为核心$GP_d(b_k)$分布提供了两个采样器。第一个是拒绝采样代理。第二个是从齐次CLWE密度推导出的精确闭式采样器,并针对其自身解析形式进行了验证。使用此实现,我们运行统计不可区分性测试,涵盖权重空间和功能性黑盒比较。在一系列稀疏比$\ ho = d_{\ ext{sparse}}/D$下,我们未发现后门模型与干净模型之间存在可检测差异的证据。我们报告了构造中哪些部分易于实现,哪些部分需要论文中未详细说明的推导。我们还强调了未尝试复现的部分,包括底层格硬度归约。我们将这项工作视为对理解Goldwasser白盒CLWE核心实际可实现性的贡献,而非新的理论结果。

英文摘要

Goldwasser et al. showed that undetectable backdoors can be planted in machine learning models trained with the Random Fourier Features (RFF) algorithm, under a hardness assumption tied to the Continuous Learning With Errors (CLWE) problem. Under standard cryptographic assumptions, even a full white-box audit of a model's weights cannot detect this class of backdoor. The construction is stated in terms of cryptographic reductions and probabilistic lemmas, without a reference implementation, and relies on secondary machinery such as the Sparse Gaussian Pancakes distribution and a homogeneous CLWE conditional density. Its realizability in ordinary numerical code is not obvious from the paper alone. This paper implements the white-box CLWE-RFF backdoor construction end to end using only numpy and scipy, to test whether this threat is realizable with commodity scientific-computing tools or requires specialized cryptographic infrastructure. We give two samplers for the core $GP_d(b_k)$ distribution. The first is a rejection-sampling proxy. The second is an exact closed-form sampler derived from the homogeneous CLWE density and verified against its own analytic form. Using this implementation, we run statistical indistinguishability tests, covering both weight-space and functional black-box comparisons. We find no evidence of detectable difference between backdoored and clean models across a range of sparsity ratios $ρ= d_{\text{sparse}}/D$. We report which parts of the construction were straightforward to realize, which required derivation not spelled out in the paper. We also highlight which parts we did not attempt to reproduce, including the underlying lattice hardness reduction. We see this work as a contribution to understanding the practical realizability of the Goldwasser white-box CLWE core, not as a new theoretical result.

发表机构

  • Laboratory for Advanced Cybersecurity Research(高级网络安全研究实验室)
  • Old Dominion University(奥多明尼昂大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑