GEO-Flag:检测与测量经GEO优化的网页内容
GEO-Flag: Detecting and Measuring GEO-Optimized Web Content
浏览论文内容
中文总结 AI 辅助
本研究针对生成式引擎优化(GEO)网页检测不足的问题,构建GEOFlagBench基准,提出干预配对训练(IPT)方法,开发GEO门控智能体系统,还部署流程测量出GEO总体流行率为8.90%。
中文摘要 AI 辅助
生成式引擎优化(Generative Engine Optimization, GEO)通过修改网页内容,提升其被生成式搜索引擎选中并引用的概率。这会让经过策略性优化的页面获得与其权威性或相关性不相称的曝光度,甚至让薄弱或虚假信息显得有充分支撑。与传统搜索不同,生成式搜索引擎会将信息合成为直接答案而非呈现竞争来源,这会进一步放大上述风险,因为评估来源的出处与权威性需要用户额外交互。尽管存在这些担忧,检测经GEO优化网页的系统性方法仍未得到充分探索。我们推出了GEOFlagBench,这是一个包含3200个网页的基准数据集,覆盖400个查询、4个领域和8个GEO优化器家族,并利用它系统性评估现有GEO检测方法。尽管最强基线的综合F1值达到0.880,但方法层面和作者身份条件评估显示出重大缺陷,且可能依赖与作者身份相关的捷径。因此,我们提出了干预配对训练(Intervention-Paired Training, IPT),该方法通过监督检测器对GEO干预和非GEO AI润色的响应来实现;在ModernBERT上,IPT将F1值从0.862提升至0.944,最差组准确率从0.725提升至0.883。我们开发了一个GEO门控智能体(Agent)系统,用于审计检测到的GEO页面中的来源层级和引用URL的可验证性。最后,我们将完整流程部署到已发布的Google Search和基于Gemini的检索结果上,针对1000个真实用户查询展开评估。在10095个可用页面中,我们估计GEO的总体流行率为8.90%,在2026年修改的页面中这一比例达到16.36%。我们的研究结果为在现实搜索生态系统中系统性检测、审计和测量GEO奠定了基础。
英文摘要
Generative Engine Optimization (GEO) modifies web content to increase its likelihood of being selected and cited by generative search engines. This can give strategically optimized pages visibility disproportionate to their authority or relevance and even make weak or false information appear well supported. Unlike conventional search, generative search synthesizes information into direct answers rather than presenting competing sources, which can further amplify these risks, as assessing source provenance and authority requires additional user interaction. Despite these concerns, systematic methods for detecting GEO-optimized webpages remain underexplored. We introduce \texttt{GEOFlagBench}, a benchmark of 3,200 web content instances spanning 400 queries, four domains, and eight GEO optimizer families, and use it to systematically evaluate existing GEO detection methods. Although the strongest baseline achieves an aggregate F1 of 0.880, method-level and authorship-conditioned evaluations reveal substantial weaknesses and potential reliance on authorship-related shortcuts. We therefore propose \emph{Intervention-Paired Training} (IPT), which supervises detector responses to GEO interventions and non-GEO AI polishing; on ModernBERT, IPT improves F1 from 0.862 to 0.944 and worst-group accuracy from 0.725 to 0.883. We develop a GEO-gated Agent system for auditing the Source Tier and verifiability of Citation URLs in detected GEO pages. Finally, we deploy the complete pipeline on released Google Search and Gemini-grounded retrieval results for 1,000 real-user queries. Across 10,095 available pages, we estimate an overall GEO prevalence of 8.90\%, reaching 16.36\% among pages modified in 2026. Our results establish a foundation for systematically detecting, auditing, and measuring GEO in real-world search ecosystems.
发表机构
- CISPA Helmholtz Center for Information Security(CISPA亥姆霍兹信息安全中心)
- HPE(慧与科技)
- University of Waterloo(滑铁卢大学)
机构由 AI 辅助整理,请以论文原文为准。