arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于角色大语言模型模拟的增强假设检验

Augmented Hypothesis Testing with Persona-Based LLM Simulations

Ziyad Benomar, Aymen Al Marjani, Paul Missault, Saab Mansour

arXiv 2609.24629首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出学习增强假设检验框架,利用未知质量的预测(总体方向或个体级)减少A/B测试样本量,经角色LLM模拟验证,在四个数据集上显著降低成本且保持统计有效性。

AI 中文摘要

A/B测试需要较大的样本量、较长的时间周期和显著的成本。当机器学习模型能够提供实验结果的辅助预测时,预测质量的不确定性使得完全替代人类实验不可行,但这些预测仍可能包含有用信号。我们提出了一种用于学习增强假设检验的原则性框架,该框架利用质量未知的预测来减少样本量,同时保持统计有效性。预测在粒度上自然存在差异,从粗略的聚合信号到细粒度的个体级估计,我们的框架同时处理了这一范围的两端:(1)对于总体层面的方向性预测,即仅可获得关于处理效应符号的二元信号时,我们采用非对称检验,并在学习增强算法范式内证明了其一致性和鲁棒性界;(2)对于个体级预测,我们引入了广义PPI++(GPPI),将预测驱动推断扩展到通过高维变换处理非线性预测误差。这两种方法都能从准确预测中获益,同时对不准确或对抗性预测保持鲁棒性。我们使用基于角色的大语言模型模拟来验证我们的框架,其中配备用户角色的AI智能体预测个体行为,作为涵盖两种粒度水平的自然预测来源。在四个真实世界数据集上的实验表明,我们的方法结合基于角色的预测,在保持严格统计有效性的同时,大幅降低了实验成本。

英文摘要

A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.

CommentsWork accepted at COLM Workshop on Agent Behavior

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑