GroundedGEO:审计生成式搜索排序中的证据缺口
GroundedGEO: Auditing the Evidence Gap in Generative Search Rankings
浏览论文内容
中文总结 AI 辅助
针对生成式搜索中无支持文本的排名优势问题,提出证据配对基准与主张级重排序器GroundedGEO,通过惩罚缺乏支持的查询相关主张来缩小可识别性缺口,并揭示标签质量与文档包覆盖率两个关键限制。
中文摘要 AI 辅助
生成式搜索系统为具有重大影响的决策对产品和服务进行排序,而发布者可以低成本地使候选文本看起来具有相关性。然而,证据状态并非文本属性,而是一种主张-证据关系:仅基于文本的排序器和防御方法无法区分诚实的详细内容与捏造的细节,从而产生可识别性缺口。我们通过一个证据配对基准(50个电商查询,1,950个案例)和一个主张级重排序器GroundedGEO来审计这一缺口,该重排序器会惩罚在所提供的文档包中缺乏支持的查询相关主张。匹配的丰富变体控制格式和篇幅;文档包孪生体在固定文本处添加证明,而精简文档包则移除证明。在冻结的列表式排序器Qwen2.5-7B上,无支持的丰富变体相对于干净候选显示出显著的归一化排名提升(跨主张配置文件为+0.065至+0.092,经Holm校正),而受支持和中性对照则无此效应;该效应依赖于模型(在MiMo-v2.5上边际显著,在GLM-5.3-Flash上不存在)。在冻结的点式评分器上,预言机证据标签将无支持的丰富变体前3名比率从0.65降至0.43(洗白率从0.61降至0.39),在lambda=40时零错误抑制;文档包孪生体在不改变文本的情况下恢复了原始比率。针对370条主张的人工金标准,所有测试的自动评判者均未通过预注册的可靠性门槛,尽管最佳本地评判者在受保护臂上保留了预言机抑制的79-100%,且实测零错误抑制。此外,剥离证明覆盖率会使错误抑制增加0.307。这些诊断效应揭示了证据通道的两个限制:标签质量和文档包覆盖率。它们并不验证自动防御的有效性,且对不利的人工金标准臂的解释仍有待裁决。
英文摘要
Generative search systems rank products and services for consequential decisions, and publishers can cheaply make candidate text look relevant. Yet evidence status is not a text property but a claim-evidence relation: text-only rankers and defenses cannot separate honest detailed content from fabricated detail, creating an identifiability gap. We audit this gap with an evidence-paired benchmark (50 e-commerce queries, 1,950 cases) and a claim-level reranker, GroundedGEO, that penalizes query-relevant claims lacking support in a supplied packet. Matched rich variants control format and volume; packet twins add attestations at fixed text, while thinned packets withdraw them. On the frozen listwise ranker Qwen2.5-7B, unsupported-rich variants show significant normalized rank gain over clean candidates (+0.065 to +0.092 across claim profiles, Holm-corrected), while supported and neutral controls do not; the effect is model-dependent (marginal on MiMo-v2.5, absent on GLM-5.3-Flash). On a frozen pointwise scorer, oracle evidence labels cut the unsupported-rich top-3 rate from 0.65 to 0.43 (laundering from 0.61 to 0.39) at lambda=40 with zero false suppression; packet twins restore the original rates without changing text. Against a 370-claim human gold, all tested automatic judges fail the preregistered reliability gate, although the best local judge retains 79-100% of oracle suppression with zero measured false suppression on protected arms. Separately, stripping attestation coverage increases false suppression by 0.307. These diagnostic effects identify two limits on the evidence channel: label quality and packet coverage. They do not validate an automatic defense, and interpretation of the adverse human-gold arm remains pending adjudication.
发表机构
- Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。