面向选区划分方案的分层抽样方法
Towards stratified sampling for redistricting plans
浏览论文内容
中文总结 AI 辅助
针对选区重划方案的高维组合空间,提出通过聚类构建“字母-单词”语法来划分候选层,并用单位分解估计层质量与通量矩阵,为未来分层抽样奠定基础。
中文摘要 AI 辅助
快速发展的算法加速了对选区重划集成(平衡图划分)的采样,然而,由于相空间的高维和组合特性,评估稀有事件和采样复杂目标度量仍然是一个核心挑战。我们解决了在该空间上进行分层抽样的一个先决条件:构建和诊断具有适当覆盖率和重叠度的候选层。我们通过将选区聚类为具有代表性的“字母”,并利用它们形成方案级别的“单词”,从而在观测到的方案上构建一种语法。这些单词上的单位分解给出了方案到各层的软分配,并使我们能够估计层质量和由重叠引起的通量矩阵。我们使用来自康涅狄格州的真实国会选区重划数据演示了这一计算流程,并考察了从一种目标分布学习到的层在相关分布下的表现。所得到的构建为未来在选区重划方案或平衡图划分空间上进行分层抽样奠定了基础。我们在此并未实现完整的分层抽样器;评估所提出的层是否能提高抽样效率或降低估计量方差留待未来工作。
英文摘要
Rapid algorithmic developments have accelerated the sampling of redistricting ensembles (balanced graph partitions), yet evaluating rare events and sampling complex target measures remains a core challenge due to the high-dimensional and combinatorial nature of the phase space. We address a prerequisite for stratified sampling on this space: constructing and diagnosing candidate strata with suitable coverage and overlap. We build a grammar on observed plans by clustering districts into representative ``letters'' and using them to form plan-level ``words.'' A partition of unity over these words gives a soft assignment of plans to strata and allows us to estimate stratum masses and an overlap-induced flux matrix. We demonstrate this computational pipeline using real-world congressional redistricting data from Connecticut and examine how strata learned from one target distribution behave under related distributions. The resulting construction provides a foundation for future stratified sampling on spaces of redistricting plans or balanced graph partitions. We do not implement a complete stratified sampler here; evaluating whether the proposed strata improve sampling efficiency or reduce estimator variance is left for future work.
发表机构
- Yale University(耶鲁大学)
- Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。