arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

识别模型质量对用户参与度的影响:一种结合合成数据验证的版本内因果估计器

Identifying Model Quality Effects on User Engagement: A Within-Version Causal Estimator with Synthetic Data Validation

John Tribbia

arXiv 2608.17187首次发表:更新:

AI 中文总结

针对LLM模型更新与用户参与度的因果识别难题,提出结合合成数据验证的版本内因果估计器,经衰减调整后可准确估计模型质量对参与度的影响。

AI 中文摘要

每个构建大语言模型(LLMs)的团队都面临一个核心挑战:离线基准显示性能提升,部署后用户参与度也会上升,但很难将因果关系与其他因素区分开。同期的营销活动、媒体报道和季节性需求会掩盖模型更新是否真正推动了参与度提升。本文提出一种新颖的因果估计方法,利用单个模型版本内不同能力的非均匀质量改进。由于能力提升不均衡(例如,编码能力大幅提升,而写作能力提升幅度较小),用户的体验质量会因任务分布不同而存在差异。这种体验质量的差异为估计提供了因果信号。我们在具有已知真实值的合成数据上验证了该方法。原始估计器可恢复81%至88%的真实效应,剩余部分因使用量估计的测量误差而损失。应用变量误差衰减调整(使用重测信度比率平均相关性为0.861)后,估计值修正为1.017(自举95%置信区间:0.915至1.109)。相比之下,朴素方法表现显著不佳,在无版本控制时仅恢复63%,使用实时而非冻结使用模式时仅恢复62%。置换检验证实该框架可区分真实因果效应与噪声。敏感性分析表明,随着前期时长从3周增加到7周,恢复率从78%单调提升至87%,凸显了样本时长与估计器精度之间的明确权衡。种子敏感性测试进一步确认了随机抽样下的稳定性。

英文摘要

Every team building Large Language Models (LLMs) faces a core challenge: offline benchmarks show performance gains and user engagement rises post deployment, but isolating cause from effect remains difficult. Simultaneous marketing, media coverage, and seasonal demand obscure whether model updates truly drive engagement gains. This paper presents a novel causal estimation approach that leverages non uniform quality improvements across capabilities within a single model version. Because capabilities improve unevenly (e.g., strong gains in coding versus modest gains in writing), users experience varied quality depending on their task distribution. This variation in experienced quality provides causal signal for estimation. We validate this approach on synthetic data with known ground truth. The raw estimator recovers 81% to 88% of the true effect, with the remainder lost to measurement error in usage estimates. Applying an errors in variables disattenuation adjustment (using a test retest reliability ratio mean correlation of 0.861) corrects the estimate to 1.017 (bootstrapped 95% CI: 0.915 to 1.109). By contrast, naive methods fail significantly, recovering only 63% without version controls and 62% using real time rather than frozen usage patterns. Permutation tests confirm the framework distinguishes true causal effects from noise. Sensitivity analysis indicates recovery improves monotonically from 78% to 87% as pre period length increases from 3 to 7 weeks, highlighting a clear tradeoff between sample duration and estimator precision. Seed sensitivity tests further confirm stability across random draws.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑