零售智能中的选择偏差校正
Selection Bias Correction in Retail Intelligence
浏览论文内容
中文总结 AI 辅助
本研究针对零售智能中因忽略长尾商品产生的选择偏差,通过模拟实验对比逆概率加权(IPW)与分层法,发现分层法在多数长尾场景下更具稳健性,为零售通胀估算提供了更优的校正方案。
中文摘要 AI 辅助
零售智能通常依赖于对热门、高周转产品的监测,可能因忽略小众商品的“长尾”而导致经济指标出现偏差。本模拟研究调查了通胀估算中的选择偏差,并在不同的数据生成过程中对比了多种校正方法。我们通过4种场景(对齐阶跃函数、平滑梯度、错位断点、多项式关系)下的400次蒙特卡洛重复实验,测试了5种设定的逆概率加权(IPW)与不同分层数的分层法的稳健性。研究发现了零售长尾场景中加权方法的基本局限:分层法在4种场景中的3种表现更优,即使在边界刻意与总体断点错位时,仍能保持低于0.04个百分点的中位误差(较IPW有116倍优势);但在平滑多项式关系场景下,带有样片倾向评分模型的IPW胜出(中位误差0.007个百分点,对比分层法的0.013个百分点),体现了方法的情境依赖性。关键的是,即使拥有完美结构知识的理想IPW设定,在阶跃函数场景下仍会产生6.06个百分点的误差,而分层法仅为0.008个百分点,这反映了正性假设(因果推断的基本要求)被违反,而非IPW方法本身的缺陷。当选择概率差异极大(90% vs 1%)时,加权方法超出了其理论设计范围。这些结果表明,在存在严重正性违反的零售长尾分布中,分层法是更稳妥的工程选择。
英文摘要
Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the "long tail" of niche items. This simulation study investigates selection bias in inflation estimation and compares correction methods across diverse data-generating processes. Through 400 Monte Carlo replications spanning four scenarios--aligned step functions, smooth gradients, misaligned breaks, and polynomial relationships--we test the robustness of Inverse Probability Weighting (IPW) with five specifications against stratification with varying strata counts. Our findings reveal fundamental limits of weighting methods in retail long-tail contexts: stratification achieves superior performance in three of four scenarios, maintaining sub-0.04pp median error even when boundaries deliberately misalign with population breaks (116x advantage over IPW). However, IPW with spline propensity models wins under smooth polynomial relationships (median error 0.007pp vs. 0.013pp), demonstrating context-dependency. Critically, even an oracle IPW specification with perfect structural knowledge achieves 6.06pp error compared to stratification's 0.008pp in step-function scenarios. This reflects violation of the Positivity Assumption--a fundamental causal inference requirement--rather than IPW methodological inferiority. When selection probabilities differ dramatically (90% vs. 1%), weighting methods operate outside their theoretical design envelope. These results demonstrate that stratification provides a safer engineering choice in retail long-tail distributions with severe positivity violations.
发表机构
- Walmart Inc.(沃尔玛公司)
机构由 AI 辅助整理,请以论文原文为准。