发表机构
Zhejiang Normal University; Zhejiang University; China Electric Power Research Institute(浙江师范大学; 浙江大学; 中国电力科学研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在低查询预算硬标签场景下生成高质量对抗文本的难题,提出基于采样的LBA方法,整合先验与后验知识构建近似分布用于采样,实验证明该方法在多语言模型上显著优于基线,生成的对抗文本语义保留性和可理解性更佳。
AI 中文摘要
在硬标签场景中,以低查询预算生成高质量对抗文本仍然是一个具有挑战性的问题。大多数现有方法依赖贪婪算法,在文本中选择一个位置进行替换,然后再替换其他位置。这种局部搜索方法可能无法发现高质量的对抗示例,并且常常导致过高的查询成本。理想情况下,最优对抗样本会考虑文本中所有可能的位置组合,但穷举搜索在计算上不切实际。为应对这一挑战,我们提出了一种基于采样的方法LBA,它通过整合先验和后验知识来构建高质量对抗示例的近似分布,并利用此分布进行采样。随着采样的进行,后验知识更新近似分布,进而指导更有效的采样。在四个数据集上对六种从小规模到大规模架构的语言模型进行的大量实验表明,LBA在所有评估指标上均显著优于现有基线。此外,基于大语言模型的评估表明,LBA生成的对抗文本在语义上更具保留性且更易理解。
英文摘要
Generating high-quality adversarial texts with low query budgets remains a challenging problem in the hard-label scenario. Most existing approaches rely on greedy algorithms, where one position in the text is selected for substitution, followed by the substitutions of other positions. This local search approach may fail to discover high-quality adversarial examples and often leads to excessive query costs. Ideally, an optimal adversarial sample would consider all possible position combinations in the text, but exhaustive search is computationally impractical. To address this challenge, we propose a sampling-based method called LBA, which constructs an approximate distribution of high-quality adversarial examples by integrating both prior and posterior knowledge, and utilizes this distribution for sampling. As sampling progresses, posterior knowledge updates the approximate distribution, which in turn guides more effective sampling. Extensive experiments on six language models, ranging from small-scale to large-scale architectures across four datasets, demonstrate that LBA significantly outperforms state-of-the-art baselines on all evaluation metrics. Additionally, LLM-based assessment indicates that LBA generates more semantically preserved and comprehensible adversarial texts.