发表机构
The Pennsylvania State University; Information Sciences and Technology(宾夕法尼亚州立大学; 信息科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出AAST框架,通过捆绑层面联合选择合成文本解决现有作者混淆方法忽略跨文档关联的问题,实验显示其可降低账户级可链接性并保留文本质量。
AI 中文摘要
在线用户常以同一身份发布多篇文本,攻击者可构建作者画像,获取比单篇文本更多的信息。现有作者混淆方法针对每篇文档独立优化隐私,忽略了使聚合攻击危险的跨文档关联。我们提出Aggregation-Aware Synthetic Text Generation(AAST,聚合感知合成文本生成)框架,通过在捆绑层面联合选择合成文本而非孤立优化每篇文本,解决该问题。AAST针对归属和验证攻击,包括生成或选择阶段未观察到攻击者参考文本所属体裁的跨体裁场景。在同体裁、跨体裁、神经及独立非神经风格度量攻击的实验表明,随着捆绑规模增大,AAST可降低账户级可链接性,同时保留语义质量、语言可接受性和情感一致性。
英文摘要
Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single text. Existing authorship obfuscation methods optimize privacy independently for each document, leaving them blind to cross-document correlations that make aggregation dangerous. We propose Aggregation-Aware Synthetic Text Generation (AAST), a framework that addresses this gap by jointly selecting synthetic texts at the bundle level rather than optimizing each text in isolation. AAST targets attribution and verification attacks, including cross-genre settings where attacker references come from a genre not observed during generation or selection. Experiments across same-genre, cross-genre, neural, and independent non-neural stylometric attacks show that AAST lowers account-level linkability as bundle size grows, while preserving semantic quality, linguistic acceptability, and sentiment alignment.
CommentsProceedings of the 2026 Conference on Empirical Methods in Natural Language Processing
Journal refThe 2026 Conference on Empirical Methods in Natural Language Processing