arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

无引擎评分:端到端验证生成引擎的确定性、抗操纵内容评分

Scoring Without the Engine: Validating a Deterministic, Manipulation-Resistant Content Score for Generative Engines, End to End

Elisha Bajemon, Andre-Louis Rochet

arXiv 2609.07559首次发表:更新:

发表机构

TW3 Partners; Citead(TW3合伙人机构; Citead)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于对抗性证伪门的协议,验证廉价确定性代理,并在生成引擎优化上端到端演示,发现因果锚点过期,重新校准后评分仅保留门强制的响应面,可测的抗操纵性、检测局限及查询条件化引用信号均被界定。

AI 中文摘要

如何验证一个廉价、确定性的代理,其对应的神谕(oracle)昂贵、限速且非平稳?我们提出了一种基于对抗性证伪门(阴性对照、剂量反应、有界放大、重复惩罚、长度中性)的协议,这些门定义并选择代理,在训练集上拟合并在留出集上确认;围绕这些门,协议界定了代理永远无法解析的内容,并在当前神谕上重新测量外部因果证据,而非假设其不变。我们在生成引擎优化(Generative Engine Optimization)上端到端地演示了该协议,其中代理是一个确定性内容评分,而该领域恰好有一项步骤失败,正如协议设计用于检测的那样:在十个现代引擎家族上重新测量唯一已发表的因果锚点(2023年效应量)显示,它们的杠杆对引用没有影响,因此这些锚点已过期;重新校准到接近零的现代向量会剥离评分中杠杆响应的成分。幸存下来的是门强制的响应面。这些门带来了一种可测量的属性:在一个包含500个来源的对抗性编辑基准上,放大评分校准后的杠杆最多为攻击者带来6分,且随剂量递减;单杠杆放大被证明有界,而上限和跨杠杆次可加性是与之一致的经验发现。在检测方面,网络垃圾邮件基线占主导地位,分布外攻击可逃避评分,因此可部署的过滤器将其叠加在这些基线之上。一个查询条件化的天际线界定了评分的引用信号(查询内Spearman相关系数为0.11),将查询无关的评分重新定位为质量过滤器而非引用预测器。我们披露并纠正了首次排序评估中的查询泄漏错误和一次失败的置信度标志;所有数字均可从发布的工件离线复现,且边际API成本为零。

英文摘要

How do you validate a cheap, deterministic proxy for an oracle that is expensive, rate-limited, and non-stationary? We present a protocol built on adversarial falsification gates (negative control, dose response, bounded amplification, duplication penalty, length neutrality) that define and select the proxy, fitted on a training split and confirmed held-out; around them it bounds what the proxy can never resolve, and re-measures external causal evidence on the current oracle rather than assuming it. We demonstrate it end to end on Generative Engine Optimization, where the proxy is a deterministic content score, and one step fails on that domain exactly as the protocol is built to detect: re-measuring the only published causal anchors (2023 effect sizes) on ten modern engine families shows their levers move citation on none, so the anchors are an expired external check; recalibrating to the near-zero modern vector strips the score of its lever-responsive components. What survives is the gate-enforced response surface. The gates buy a measured property: on a 500-source benchmark of adversarial edits, amplifying the score's calibrated levers gains an attacker at most 6 points, and decreases with dose; single-lever amplification is provably bounded, while the cap and cross-lever sub-additivity are empirical findings consistent with it. On detection, web-spam baselines dominate and out-of-distribution attacks evade the score, so the deployable filter layers it over them. A query-conditioned skyline bounds the score's citation signal (within-query Spearman 0.11), repositioning query-agnostic scores as quality filters rather than citation predictors. A query-leakage bug in our first ranking evaluation and a failed confidence flag are disclosed and corrected; every number reproduces offline from released artifacts at zero marginal API cost.

Comments42 pages, 4 figures. Code and data: https://github.com/TW3-Partners-OS/geo-score-reproducibility & More informations on Citead.com

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑