发表机构
Graduate School of Information Science and Technology, The University of Tokyo; Algorithms Group, IT University of Copenhagen(东京大学情报理工学研究科; 哥本哈根信息技术大学算法组)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对最优分段线性近似构建的PGM-index提出PGM-attack中毒攻击,仅10%密钥中毒可使分段数增120倍,攻击可迁移至其他学习型索引,揭示其优化目标存在固有漏洞。
AI 中文摘要
PGM-index(Ferragina与Vinciguerra,VLDB'20)是最实用的学习型索引之一,因其理论简洁性和持续优异的实验性能而著称,它基于使分段数量最小的最优分段线性近似(PLAs)构建。本文研究该最优PLA本身对中毒攻击的敏感性,提出一种高效的中毒攻击方法PGM-attack,其通过顺序插入对抗性密钥来增加最终的分段数量,同时开发了一种方法可推导在任意插入下能达到的分段数量的理论上界。实验表明,仅对10%的密钥进行中毒,PGM-attack就能使分段数量最多增加120倍;在所有评估实例中,本文提出的依赖实例的上界最多为PGM-attack达到的分段数量的1.92倍,证明PGM-attack至少达到了最优值的52%。分段数量的增加会使PGM-index的规模最多扩大120倍,此外,该攻击还能迁移到其他学习型索引,尤其是会大幅增加基于PLA的索引的规模。研究结果表明,尽管PGM-index的PLA具有最优性,但其优化目标存在固有漏洞,这为未来学习型索引的鲁棒感知目标设计提供了动机,本文代码可在指定URL公开获取。
英文摘要
The PGM-index (Ferragina and Vinciguerra, VLDB'20) is one of the most practical learned indexes, owing to its theoretical elegance and consistently strong empirical performance. It is built on optimal piecewise linear approximations (PLAs) that minimize the number of segments. In this paper, we ask how sensitive this optimal PLA itself is to poisoning attacks. We propose PGM-attack, an efficient poisoning attack that sequentially inserts adversarial keys to inflate the resulting number of segments, and we develop a method for deriving theoretical upper bounds on the number of segments attainable under arbitrary insertions. Our experiments show that poisoning only 10% of the keys allows PGM-attack to increase the segment count by up to 120x. On every evaluated instance, our instance-dependent upper bound is at most 1.92x the segment count attained by PGM-attack, certifying that PGM-attack achieves at least 52% of the optimum. This increase in the number of segments enlarges the PGM-index by up to 120x. Moreover, the attack also transfers to other learned indexes, substantially inflating the index size of PLA-based ones in particular. Our results reveal that, despite the optimality of its PLAs, the PGM-index has an intrinsic vulnerability rooted in its optimization objective, motivating robustness-aware objective design for future learned indexes. Our code is publicly available at https://github.com/atsukisato/pgm-attack.