发表机构
The Hong Kong Polytechnic University; Commonwealth Scientific and Industrial Research Organisation (CSIRO)(香港理工大学; 联邦科学与工业研究组织(CSIRO))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TraceGuard提出基于跨特征排名一致性的自适应过滤方法,利用六个语料级特征识别投毒样本,在19种攻击配置下平均移除98.4%投毒示例,残余攻击指标低至1%。
AI 中文摘要
多模态训练依赖于从外部来源收集的图像-文本语料库,这为攻击者提供了投毒数据的机会。隐蔽攻击可以保留看似合理的图像-文本对,同时隐藏检测器所使用的差异,因此表面上干净的数据仍可能使训练后的模型发生偏移。因此,我们提出疑问:一个投毒集必须保留哪些属性才能使攻击保持有效。一个小的投毒集在训练期间仍必须施加足够的集体影响,以诱发攻击者的目标行为。我们根据攻击模式出现的频率以及携带该模式的示例对模型的联合影响强度来分析这种影响。这一分析促使我们提出六个语料库级特征,这些特征在不训练受害者模型的情况下检查跨模态邻域、重复文本以及文本跨度擦除后的变化。我们引入了TraceGuard,一种自适应的基于排名的过滤方法,它利用互补特征排名之间的一致性来识别可疑示例。它通过共享模式细化所选集合,并在不知道攻击或投毒率的情况下为每个语料库自适应调整移除阈值。在涵盖图像-文本学习、生成式视觉-语言模型微调和编码器-迁移测试的19种攻击配置中,TraceGuard平均移除了98.4%的投毒示例和5.4%的干净示例。在过滤后的语料库上训练后,13种配置中的残余攻击指标最多为1%。匹配移除对照和消融实验支持了样本选择和自适应移除的贡献。压力测试还识别了自适应攻击下的检测失败以及对无投毒语料库的不必要移除。
英文摘要
Multimodal training relies on image-text corpora collected from external sources, creating opportunities for attackers to poison the data. Stealthy attacks can preserve plausible image-text pairs while concealing the differences used by detectors, so apparently clean data can still redirect the trained model. We therefore ask which properties a poison set must preserve for the attack to remain effective. A small poison set must still exert enough collective influence during training to induce the attacker's target behavior. We analyze this influence in terms of how often an attack pattern occurs and how strongly the examples carrying it jointly affect the model. This analysis motivates six corpus-level features that examine cross-modal neighborhoods, recurring text, and changes after text-span erasure without training the victim model. We introduce TraceGuard, an adaptive rank-based filtering method that uses agreement among complementary feature rankings to identify suspicious examples. It refines the selected set through shared patterns and adapts the removal threshold to each corpus without knowing the attack or poison rate. Across 19 attack configurations spanning image-text learning, generative vision-language model fine-tuning, and encoder-transfer tests, TraceGuard removes an average of 98.4% of poisoned examples and 5.4% of clean examples. After training on the filtered corpora, the residual attack metric is at most 1% in 13 configurations. Matched-removal controls and ablations support the contributions of sample selection and adaptive removal. Stress tests also identify detection failures under adaptive attacks and unnecessary removal on poison-free corpora.
Comments42 pages