arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

响应边际遗漏之处:功能后门恢复的自适应查询复杂度

What Response Marginals Miss: Adaptive Query Complexity of Functional Backdoor Recovery

Yunjae Hwang, Byoungjin Seok

arXiv 2610.07771首次发表:更新:

发表机构

Korea University; Hansung University(高丽大学; 汉城大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过构造反例证明,在标签反馈下,攻击成功率与标签分布不足以决定功能后门恢复的自适应查询复杂度,其最优复杂度可呈对数与线性差异,并实验验证。

AI 中文摘要

功能后门恢复旨在寻找任何攻击成功率至少达到给定阈值的触发器,而非恢复被植入的触发器。我们研究了在标签反馈(返回预测类别标签)下完成此任务所需的最小查询数量。我们构造了两个有限的受害模型族,它们对于每个受害者和触发器候选都具有完全相同的攻击成功率。在均匀随机选择的受害者下,每个查询返回标签的分布也相同。这些族构成了一个明确的反例:尽管这些量匹配,它们的最优自适应查询复杂度分别为\\(\Theta(\log H)\\)和\\(\Theta(H)\\),其中\\(H\\)是可能受害者的数量。差异的产生是因为相同的非目标标签与不同的受害者集合相关联,因此连续的查询以不同的速率消除可能的受害者。当响应被简化为仅报告是否返回目标标签的二元反馈时,这种分离消失。对于低于一的每个固定失败概率,这种分离也持续存在。最后,我们用训练好的CIFAR-10 ResNet-18分类器实现了相同的恢复问题,并验证了预测的最优查询预算。这些结果表明,攻击成功率和每个查询返回标签的分布不足以确定功能后门恢复的查询复杂度。

英文摘要

Functional backdoor recovery finds any trigger whose attack success rate is at least a given threshold rather than to recover the planted trigger. We study the minimum number of queries required for this task under label feedback which returns the predicted class label. We construct two finite families of victim models that have exactly the same attack success rate for every victim and trigger candidate. The distribution of returned labels for every query is also identical under a uniformly chosen victim. These families form an explicit counterexample that despite the matched quantities, their optimal adaptive query complexities are \(Θ(\log H)\) and \(Θ(H)\) where \(H\) is the number of possible victims. The difference arises because the same non-target labels are associated with different sets of victims, so successive queries eliminate possible victims at different rates. This separation disappears when the response is reduced to binary feedback, which reports only whether the target label is returned. The separation also persists for every fixed failure probability below one. Finally, we realize the same recovery problems with trained CIFAR-10 ResNet-18 classifiers and verify the predicted optimal query budgets. These results show that attack success rate and the distribution of returned labels for each query are insufficient to determine the query complexity of functional backdoor recovery.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑