当优化变为操纵:防御生成式搜索免受恶意生成式引擎优化攻击
When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization
浏览论文内容
中文总结 AI 辅助
本文提出GEO Defender两阶段防御方案,无需微调目标LLM,可将GEO攻击平均成功率从50.32%降至6.20%,保留94.12%良性证据使用,能泛化到未见过的GEO攻击。
中文摘要 AI 辅助
本文聚焦于防御生成式搜索引擎抵御恶意生成式引擎优化(Generative Engine Optimization,GEO)攻击,这类攻击会改写网页文档以匹配搜索引擎的引用偏好,进而操纵生成的答案。近期的GEO方法已从人工改写发展为自动化及智能体优化,大幅提升了目标文档在生成答案中的可见性。然而,防御此类操纵面临两大核心挑战:攻击文档与原始内容在事实上保持一致,致使事实验证和困惑度过滤失效;且它们放大的特征也同样是高质量良性内容的特征。为应对这些局限,我们提出GEO Defender,这是一种与攻击链对齐的两阶段防御方案,无需对目标大语言模型(Large Language Model,LLM)进行微调。GEO Defender由Shield重排器(Shield Reranker)和无训练的Shield生成(Training-Free Shield Generation,TFSG)构成。具体而言,Shield重排器在冻结的基础重排器上学习基于偏好的防御残差,降低经GEO改写的文档的排名,同时保留相关性判断;TFSG将防御结果提炼为自然语言经验库,指导目标LLM在推理阶段的源使用。针对两个最先进的闭源LLM和三个开源LLM,在七种GEO攻击上开展的实验表明,GEO Defender将平均攻击成功率从50.32%降至6.20%,保留了94.12%的良性证据使用,维持了答案质量,且能泛化到构造实例之外的未见过的攻击。
英文摘要
This paper focuses on defending generative search engines against malicious Generative Engine Optimization (GEO), which rewrites web documents to match engines' citation preferences and thereby manipulates generated answers. Recent GEO methods have advanced from hand-crafted rewriting to automated and agentic optimization, substantially increasing the visibility of target documents in generated answers. However, defending against such manipulation poses two major challenges: attack documents remain factually consistent with their originals, rendering fact verification and perplexity filtering ineffective, and the features they amplify equally characterize high-quality benign content. To address these limitations, we propose GEO Defender, a two-stage defense aligned with the attack chain that requires no fine-tuning of the target LLM. GEO Defender consists of Shield Reranker and Training-Free Shield Generation (TFSG). Specifically, Shield Reranker learns a preference-based defensive residual over a frozen base reranker, demoting GEO-rewritten documents while preserving relevance judgments, and TFSG distills defense outcomes into a natural-language experience library that guides the target LLM's source use at inference. Experiments on two state-of-the-art closed-source LLMs and three open-source LLMs across seven GEO attacks demonstrate that GEO Defender reduces the average attack success rate from 50.32% to 6.20%, retains 94.12% of benign-evidence use, preserves answer quality, and generalizes to unseen attacks from construction instances.
发表机构
- University of the Chinese Academy of Sciences(中国科学院大学)
- Zhongguancun Laboratory(中关村实验室)
机构由 AI 辅助整理,请以论文原文为准。