发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出基于数据驱动角色的LLM智能体模拟A/B测试,通过研究多方面因素,在40项A/B测试基准上达到0.75-0.90方向准确率,为低成本实验预筛选提供可行路径。
AI 中文摘要
A/B测试是评估产品变更的黄金标准,但每项实验都需要真实用户流量、工程投入及数周的测量时间。我们提出一种模拟框架,该框架使用基于真实用户行为信号的数据驱动角色(首次出现时补充:即对用户特征的刻画)的大语言模型(LLM)驱动智能体来预测A/B测试结果。与依赖合成或基于规则的角色的现有研究不同,我们的智能体由匿名化行为数据构建,这些数据包括活动模式、参与信号及推断出的人口统计信息,从而实现更忠实的总体建模。我们将A/B测试模拟视为结构化问答任务,并系统研究了四个方面:(i)问题设计格式;(ii)角色数据源及领域对齐的影响;(iii)每个角色的行为深度与总体多样性之间的权衡;(iv)高效的总体子采样。在涵盖两种指标类型的40项A/B测试基准上,我们的最佳配置根据测试指标的不同达到了0.75-0.90的方向准确率,证明数据驱动的角色是实现快速、低成本实验预筛选的可行途径。
英文摘要
A/B testing is the gold standard for evaluating product changes, but each experiment requires real user traffic, engineering effort, and weeks of measurement. We propose a simulation framework that predicts A/B test outcomes using LLM-powered agents conditioned on data-driven personas grounded in real user behavioral signals. Unlike prior work that relies on synthetic or rule-based personas, our agents are constructed from anonymized behavioral data-activity patterns, engagement signals, and inferred demographics-enabling more faithful population modeling. We frame A/B test simulation as a structured question task and systematically study (i) question design formats, (ii) the impact of persona data source and domain alignment, (iii) the trade-off between per-persona behavioral depth and population diversity, and (iv) efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, our best configuration achieves 0.75-0.90 directional accuracy depending on the test metric, demonstrating that data-driven personas are a viable path toward fast, low-cost experiment pre-screening.
CommentsWork accepted at EMNLP 2026 Industry Track