arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25010cs.AIcs.CLcs.CY

合成人物能否预测真实受众反应?一项模拟到真实的研究:无人物基线优于基于人物的文案模拟

Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation

Alexandre Cristovão Maiorano

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过模拟到真实实验发现,在预测标题点击率时,无人物LLM基线优于基于真实受众人口统计的十人物面板,表明合成人物模拟会引入偏差,不如直接询问模型。

中文摘要 AI 辅助

营销人员越来越多地使用大型语言模型(LLM)作为“合成人物”来预测文案发布前受众的反应,这一做法受到以下证据的鼓励:基于画像条件的LLM能够模拟人类样本。但这一预测对真实行为是否有效——以及人物机制是否真的有帮助?我们利用Upworthy研究档案(Upworthy Research Archive)进行了一项模拟到真实的有效性研究,该档案包含数千个在共享真实流量上进行的标题A/B测试,并记录了点击率,作为留出的真实基准。我们将一个基于真实受众人口统计特征的十人物小组与一个无人物零样本基线进行比较,后者仅询问模型典型读者点击的可能性。两个发现尤为突出。首先,真实基准的可靠性是约束条件:大多数A/B测试没有统计学上可区分的胜者,因此有效性只能在可靠子集(n=399)上衡量。其次,与人物模拟的前提相反,人物条件化降低了预测有效性:无人物基线对变体的排序明显更好(Kendall τ=0.361,中等效应;top-1准确率49.2%),优于人物小组(τ=0.084;top-1 34.6%),且置信区间不重叠。直接询问模型能够获得准确的总体水平先验;而强迫模型扮演特定人物则会引入偏差和噪声。该结果在三个独立的Upworthy数据划分中重复出现,在不同领域的新闻数据集上方向一致,并且对随机种子、提示措辞和模型选择具有鲁棒性——跨越三个Gemini层级和不同的模型家族(OpenAI gpt-4.1,显著配对差异)。结论是:对于预测总体参与度,普通LLM排序器优于人物模拟——合成人物不仅预测能力弱,而且比不使用它们更差。所有数字均来自公开的、以工件为先的复现包。

英文摘要

Marketers increasingly use large language models (LLMs) as "synthetic personas" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour - and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive - thousands of headline A/B tests on shared real traffic, with measured click-through - as held-out ground truth. We compare a ten-persona panel, grounded in the real audience's demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall τ = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel (τ = 0.084; top-1 34.6%), with non-overlapping confidence intervals. Asking the model directly taps an accurate population-level prior; forcing it to role-play specific personas injects bias and noise. The result replicates across three independent Upworthy splits, holds in direction on a different-domain news dataset, and is robust to seed, prompt phrasing, and model choice - across three Gemini tiers and a different model family (OpenAI gpt-4.1, significant paired gap). The takeaway: for predicting aggregate engagement, a plain LLM ranker beats persona simulation - synthetic personas are not merely a weak predictor, they are worse than not using them. All numbers regenerate from a public, artifact-first replication package.

补充信息

↑