AI 中文总结
该研究针对交错双重差分法中处理效应的异质性问题,采用狄利克雷过程混合先验提出部分同质性估计方法,可在无偏的同时提升效率,在模拟与实际应用中均表现良好。
AI 中文摘要
在交错双重差分法(DiD)设计中,个体在不同日历时间接受处理,因此处理效应并非单一数值,而是一组分队列-时间单元格的队列平均处理效应(CATTs),每个队列对应一个队列-时间单元格。标准完全灵活估计量将每个CATT作为单独参数进行估计,虽无偏但效率低下,而将所有CATT合并为单一双向固定效应(TWFE)系数虽高效,但当异质性真实存在时,会对个体效应产生偏差。我们将这两种极端选择表述为队列-时间单元格上的划分选择问题,并通过对CATTs采用狄利克雷过程(DP)混合先验来解决该问题。该模型偏好简约分组且不固定分组数量,折叠吉布斯采样器可生成点估计、边缘化未知划分的可信区间以及每对CATT的共聚类概率。在误差方差固定且对划分施加成对惩罚的情况下,最大后验(MAP)划分可简化为带ℓ₀惩罚的回归,将贝叶斯公式与同质性追踪文献关联起来。在校准模拟中,若不同效应足够可区分,该模型相比完全灵活估计量可将队列-时间效应的抽样方差降低26%至52%,且无合并估计量的偏差;后验通过对未知划分取平均,实现接近名义水平的置信区间覆盖率。在两项应用中,该方法在队列-时间效应真实异质的案例中恢复了可提升精度的部分同质性结构,而在队列-时间效应非异质的案例中表明完全合并是合适的。
英文摘要
In staggered difference-in-differences (DiD) designs, units enter treatment at different calendar times, so the treatment effect is not a single number but a set of Cohort-Average Treatment effects on the Treated (CATTs), one per cohort-time cell. Estimating every CATT as its own parameter, as the standard fully flexible estimator does, is unbiased but inefficient when some of these effects are in fact equal, whereas pooling them all into a single two-way fixed effects (TWFE) coefficient is efficient but, whenever the heterogeneity is genuine, biased for the individual effects. We frame the choice between these extremes as a partition-selection problem on the cohort-time cells and address it with a Dirichlet Process (DP) mixture prior on the CATTs. The model favors parsimonious groupings without fixing their number, and a collapsed Gibbs sampler delivers point estimates, credible intervals that marginalize the unknown partition, and co-clustering probabilities for every pair of CATTs. With the error variance held fixed and a pairwise penalty placed on the partition, a maximum a posteriori (MAP) partition reduces to an $\ell_0$-penalized regression, connecting the Bayesian formulation to the homogeneity-pursuit literature. In a calibrated simulation, the model cuts the sampling variance of the cohort-time effects by 26--52\% relative to the fully flexible estimator, without the pooled estimator's bias, provided the distinct effects are separated enough to be recovered, and the posterior delivers near-nominal confidence-interval coverage by averaging over the unknown partition. In two applications the method recovers a precision-improving partial-homogeneity structure in one, where the cohort-time effects are genuinely heterogeneous, and reports that full pooling is adequate in the other, where they are not.