发表机构
Squoosh(Squoosh)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究多模态LLM能否仅通过网页截图预测A/B测试结果,发现Gemini 3 Flash法官表现不佳,提出投票边际关卡隔离置信判断可提升性能,还在人类中重现了共识与真实结果脱节的现象,并发布了相关研究资源。
AI 中文摘要
多模态大语言模型能否仅通过网页截图预测网页的哪个版本会在真实A/B测试中获胜?我们报告了据我们所知最完整的答案,来自对真实转化测试进行的为期六周的预注册实验:大多不能——且例外情况可提前识别。在330项真实A/B测试中,Gemini 3 Flash法官的Cohen's kappa值为0.14,但在可信赖的(具有统计显著性的)一半标签上,证据无定论(kappa=0.11,置信区间包含0)。我们发现,领先CRO机构目录中44%的“真实值”标签来自非显著性测试,且法官与不可靠标签的一致性高于可靠标签——这是标注者与模型之间的共同先验,而非预测。所有标准改进手段(成本高2.8倍的前沿模型、提示词重新设计、刺激保真度、变更类型先验)均未通过其预注册关卡。法官的置信判断有所不同:投票边际关卡隔离出一个子集(覆盖49%),在显著性标签上达到kappa=0.31。我们直接测量该机制——不同模型或提示词的法官之间的一致性kappa值为0.74-0.88,但与真实结果的一致性仅约0.2,因此16票的小组约相当于2个有效独立投票——并在人类中重现了这一现象:15名CRO专家之间的一致性(评分者间kappa=0.53),但与真实结果的一致性仅为随机水平(kappa≈0)。共识(无论是人类还是模型的)是可重复、有说服力的,但并非证据。我们发布了我们的预注册内容、锁定关卡、负面结果、统计工具、人类响应以及一个声明账本,其中每个数字都带有证据层级。
英文摘要
Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aware of, from six weeks of pre-registered experiments on real conversion tests: mostly no -- and the exceptions are identifiable in advance. On 330 real A/B tests a Gemini 3 Flash judge reaches Cohen's kappa = 0.14, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (kappa = 0.11, CI includes zero). We show that 44% of the "ground-truth" labels in a leading CRO agency's catalog come from non-significant tests, and that the judge agrees more with the unreliable labels than the reliable ones -- a shared prior between labeler and model, not prediction. Every standard improvement lever (a 2.8x more expensive frontier model, prompt redesign, stimulus fidelity, change-type priors) fails its pre-registered gate. The judge's confident calls are different: a vote-margin gate isolates a subset (49% coverage) reaching kappa = 0.31 on significant labels. We measure the mechanism directly -- judges differing in model or prompt agree with each other at kappa = 0.74-0.88 while agreeing with real outcomes at only ~0.2, so a 16-vote panel carries about 2 effective independent votes -- and we reproduce it in humans: 15 CRO experts agree with each other (inter-rater kappa = 0.53) but score at chance against real outcomes (kappa ~ 0). Consensus, human or model, is reproducible, persuasive, and not evidence. We release our pre-registrations, locked gates, negative results, statistical harness, human responses, and a claims ledger in which every number carries an evidence tier.
Comments15 pages. Pre-registered experimental program with a public, tiered claims ledger; includes powered negative results, a label-validity audit, a cross-judge shared-prior measurement (n_eff ~ 2 of 16 votes), and a first-party 15-expert human baseline. Pre-registrations, statistical harness, human responses, and the full experiment ledger are released