发表机构
PayPal AI(贝宝人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Augur通过知识图谱和人物市场离线预演产品/政策变更反应,发现前沿与开放模型差距源于评估规格不足而非能力差异,并验证了合成反应的有效性。
AI 中文摘要
在产品或政策变更发布之前,关键问题在于人们将如何反应。Augur在离线状态下预演这种反应:它从变更文档中构建一个类型化知识图谱,填充一个基于真实人物画像的市场,模拟交互,并返回一份可审计的决策备忘录,推荐五种行动之一。我们汇编了Gold-50,包含五十个真实的产品和政策事件,其现实世界结果已知,并根据公开记录进行裁定,然后将五路发布判定与之对照评分。我们的核心发现是方法论上的且为负面的:前沿云模型与我们微调并离线服务的开放权重模型之间的大部分测量差距,归因于评估规格不足,而非能力差异。我们通过三种方式证明这一点。首先,仅提示词范围就能主导得分:在保持权重、案例和评分器不变的情况下,一个系统——基于Qwen3-32B的LoRA-SFT适配器——得分从0%摆动到73%。其次,在匹配的2x2消融实验中,仅在提示词中定义决策分类法(不改变模型)就能使每个前沿模型提升+24至+34个百分点;在规格不足的提示词下,离线服务的Qwen3-32B LoRA-SFT击败了所有三个前沿模型(配对McNemar检验,Holm校正),而一旦提示词公平,则未检测到与其中任何一个的显著差异。第三,与蒸馏教师的一致性上升但准确性并未随之提高,完整流程放大了系统性的“过度悲观”偏差,而非改善判定。另外,我们单独验证了反应层本身:跨越四个模型家族的盲审评估者发现,合成反应恢复了公众实际提出的67-90%的关切,且一项预注册消融实验定位了其价值——在决策最困难处价值最大,在接近上限处冗余。重新生成此处每个数字和图表的流程可从作者处获取。
英文摘要
Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it. Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic "over-doom" bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.
Comments19 pages, 15 figures, 11 tables