第三方网站流量估计是否保留因果变异?
Do Third-Party Web Traffic Estimates Preserve Causal Variation?
浏览论文内容
中文总结 AI 辅助
本研究探讨第三方平台(如Similarweb)的流量估计是否保留因果变异,发现其处理机制可能产生非经典测量误差,从而削弱或扭曲因果推断结果。
中文摘要 AI 辅助
研究人员在无法获得第一方分析数据时,日益依赖Similarweb和Semrush等第三方平台来测量网站流量。然而,这些平台报告的是模型生成的估计值而非原始数据,这引发了对其测量是否保留因果推断所需的时间与跨来源变异的问题。作为激励性诊断,我们考察了两次有记录的搜索引擎中断事件中报告的引荐流量;可见断点的缺失说明了为何不能理所当然地认为识别变异得到保留。随后,我们刻画了三种机制——来源内平滑、跨来源泄漏以及处理引发的校准误差——平台处理通过这些机制可能产生非经典的结果测量误差。分析结果和一个风格化的双重差分模拟表明,这种误差可能削弱、放大甚至逆转估计的处理效应。我们的研究结果警示,不应将基于黑盒专有模型的第三方流量测量用于因果推断。
英文摘要
Researchers increasingly rely on third-party platforms such as Similarweb and Semrush to measure web traffic when first-party analytics are unavailable. Yet these platforms report model-generated estimates rather than raw data, raising questions about whether their measures preserve the temporal and cross-source variation required for causal inference. As a motivating diagnostic, we examine reported referral traffic around two documented search-engine outages; the absence of visible discontinuities illustrates why preservation of identifying variation cannot be taken for granted. We then characterize three mechanisms, within-source smoothing, cross-source leakage, and treatment-induced calibration error, through which platform processing can generate nonclassical outcome measurement error. Analytical results and a stylized difference-in-differences simulation show that this error can attenuate, amplify, or reverse estimated treatment effects. Our findings caution against using third-party traffic measures based on black-box proprietary models for causal inference.
发表机构
- University of Washington(华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。