Lilith:训练-推理触发器偏移下的后门泛化
Lilith: Backdoor Generalization under Training-Inference Trigger Shift
浏览论文内容
中文总结 AI 辅助
Lilith是一种黑盒锚点到族框架,可实现训练-推理触发器偏移下的后门泛化,在跨多维度实验中展现高攻击成功率且效用退化有限,揭示了精确触发器评估的遗漏威胁。
中文摘要 AI 辅助
机器学习服务日益依赖公共数据、第三方提供商和外包训练,为数据投毒攻击创造了机会——这类攻击在植入持续恶意行为的同时保留良性效用。然而,现有后门研究大多评估精确触发器复用、训练阶段暴露的触发器多样性,或沿预定义变换轴的变体,因此存在关键盲点:从某一训练触发器学到的后门能否泛化到受害者训练中不存在的推理触发器族?我们将此问题表述为训练-推理触发器偏移下的后门泛化,并提出Lilith,一种黑盒锚点到族框架。仅使用不相交的代理资源,Lilith先通过单个训练锚点诱导紧凑的目标侧漏洞,再构建保留锚点诱导表示几何的有界仅推理族。我们通过锚点清除和族可达性表征该机制,推导局部正则性和有界代理-受害者差异下的全族目标保持充分条件。在跨数据集、架构、投毒率和防御措施的实验中,Lilith实现了高全族攻击成功率,且效用退化有限、触发器泛化差距小。额外分析表明,族激活依赖表示对齐而非提案机制,揭示了精确触发器评估所忽视的更广泛威胁。
英文摘要
Machine-learning services increasingly rely on public data, third-party providers, and outsourced training, creating opportunities for data-poisoning attacks that implant persistent malicious behavior while preserving benign utility. However, existing backdoor studies largely evaluate exact trigger reuse, training-exposed trigger diversity, or variations along predefined transformation axes. They therefore leave a critical blind spot: whether a backdoor learned from one training-time trigger can generalize to an inference-time trigger family absent from victim training. We formulate this problem as backdoor generalization under training--inference trigger shift and introduce Lilith, a black-box anchor-to-family framework. Using only disjoint surrogate resources, Lilith first induces a compact target-side vulnerability with a single training anchor, then constructs a bounded inference-only family that preserves the anchor-induced representation geometry. We characterize this mechanism through anchor clearance and family reach, deriving sufficient conditions for family-wise target preservation under local regularity and bounded surrogate--victim discrepancy. Experiments across datasets, architectures, poisoning rates, and defenses show that Lilith achieves high family-wise attack success with limited utility degradation and a small trigger generalization gap. Additional analyses show that family activation depends on representation alignment rather than the proposal mechanism, exposing a broader threat overlooked by exact-trigger evaluation.
发表机构
- Zhejiang University(浙江大学)
- College of Computer Science and Technology(计算机科学与技术学院)
- School of Software Technology(软件学院)
- Chongqing University(重庆大学)
- School of Big Data & Software Engineering(大数据与软件工程学院)
- Qilu University of Technology (Shandong Academy of Science)(齐鲁工业大学(山东省科学院))
机构由 AI 辅助整理,请以论文原文为准。