AI智能体能模拟A/B测试结果吗?智能体实验的验证框架
Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation
浏览论文内容
中文总结 AI 辅助
该研究提出智能体实验的验证框架S-RCT,用AI智能体模拟A/B测试结果,经67个营销A/B测试验证,校准后可降低预测误差,为实验者提供智能体信号支持。
中文摘要 AI 辅助
A/B测试仍是科技行业推出新功能的标准方式,不过每项实验都会消耗真实流量、工程精力,且需数周时间。我们将这个问题形式化为模拟随机对照试验(S-RCT),推导了两层误差分解,将智能体近似误差与 subsampling 误差分离,以便针对性改进各部分。该框架与智能体无关:任何行为模型,无论是微调的专用模型还是通用基础模型,都可作为模拟引擎。在67个历史营销A/B测试上验证后,使用现成基础模型的基线S-RCT能捕捉方向信号(符号重叠度0.70),但会系统性高估效应幅度。两阶段预期间校准协议使平方预测误差(去除不可约测量噪声后)降低约77倍;被试内设计(每个智能体同时接触两组)使标准误差降低约2.4倍。我们讨论了当前方法的局限性,并确定了实验者可从智能体信号中受益的应用场景。
英文摘要
A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles and contextual descriptions of the intervention---simulate outcomes accurately enough to vet candidate treatments before committing live traffic? We formalize this question as a \emph{Simulated Randomized Controlled Trial} (S-RCT) and derive a two-layer error decomposition that separates agent approximation error from subsampling error, enabling targeted improvements to each. The framework is agent-agnostic: any behavioral model---from a fine-tuned specialist to a general-purpose foundation model---can serve as the simulation engine. Validated on 67 historical marketing A/B tests, a baseline S-RCT using an off-the-shelf foundation model captures directional signal (sign overlap 0.70) but systematically overshoots effect magnitudes. A two-phase pre-period calibration protocol reduces the squared prediction error (after removing irreducible measurement noise) by ${\sim}77\times$; a within-subject design---where each agent is exposed to both arms---reduces standard errors by ${\sim}2.4\times$. We discuss limitations of the current approach and identify applications where experimenters stand to benefit from agentic signals.
发表机构
- Amazon(亚马逊公司)
机构由 AI 辅助整理,请以论文原文为准。