arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于对感知哈希算法进行有针对性的黑盒攻击的大语言模型引导的程序进化

LLM-Guided Program Evolution for Targeted Black-Box Attacks on Perceptual Hash Algorithms

A. Krylov, D. Rakhov, V. Veselova, D. Bolokhov, Oleg Y. Rogov

arXiv 2607.11472首次发表:更新:

AI 中文总结

研究针对感知哈希算法的黑盒攻击,提出基于GigaEvo和OpenEvolve的进化框架,通过综合分数评估攻击性能,实验表明该方法在减少查询次数、降低L2失真方面优于现有基线,揭示了相关漏洞并推动鲁棒方案开发。

AI 中文摘要

感知哈希算法(PHA)广泛用于检测良性变换下的图像伪造,但它们对对抗性选择的扰动的鲁棒性仍知之甚少,且很少有可证明的保证。我们提出了一种基于GigaEvo和OpenEvolve的新颖进化框架,用于对感知哈希算法进行有针对性的第二图像攻击。我们使用综合分数评估攻击性能,该分数综合考虑了与目标哈希的归一化汉明距离低于阈值p的对抗性图像的比例(攻击成功率)、对哈希函数发出的查询数量以及相对于原始图像的L2失真。在30对ImageNet图像上对四种已部署的PHA(pHash、PDQ、PhotoDNA、NeuralHash)进行的实验表明,我们的进化方法在对哈希函数的查询次数大幅减少的情况下,实现了与现有黑盒基线相当或更好的ASR,同时产生的对抗性图像相对于原始图像具有更低的L2失真。与基于梯度的方法不同,我们的框架不需要PHA架构的内部知识,并且自然地处理哈希输出的不可微、离散性质。这些结果揭示了广泛部署的内容审核管道中以前未报告的漏洞,并推动了可证明强大的感知哈希方案的开发。

英文摘要

Perceptual hash algorithms (PHAs) are widely deployed to detect image forgery under benign transformations, yet their robustness against adversarially chosen perturbations remains poorly understood and rarely comes with provable guarantees. We propose a novel evolutionary framework based on GigaEvo and OpenEvolve for targeted second-image attacks on perceptual hash algorithms. We assess attack performance using a composite score that jointly accounts for the fraction of adversarial images whose normalized Hamming distance to the target hash falls below threshold p (Attack Success Rate), the number of queries issued to the hash function, and the L2 distortion relative to the original image. Experiments on four deployed PHAs (pHash, PDQ, PhotoDNA, NeuralHash) across 30 ImageNet image pairs demonstrate that our evolutionary approach achieves comparable or better ASR than existing black-box baselines using substantially fewer queries to the hash function, while simultaneously producing adversarial images with lower L2 distortion relative to the originals. The best evolved programs reduce the pre-defined composite attack score relative to the best optimized seed by 41.2% for NeuralHash, 38.3% for PDQ, 34.0% for pHash, and 8.1% for PhotoDNA. Unlike gradient-based methods, our framework requires no internal knowledge of PHA architectures and naturally handles the non-differentiable, discretized nature of hash outputs. These results reveal previously unreported vulnerabilities in widely deployed content-moderation pipelines and motivate the development of provably robust perceptual hashing 1schemes.

Comments15 pages, 2 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑