复杂调查中可区分的类别分组:设计能区分哪些类别?
Distinguishable Category Groupings in Complex Surveys: Which Categories Can the Design Tell Apart?
浏览论文内容
中文总结 AI 辅助
针对调查中类别比较的可靠性,提出两种方法识别可区分分组,并在机动车碰撞原因调查中验证,发现多数分组不可区分。
中文摘要 AI 辅助
调查通常估计多个类别中各类别所占的比例,例如事件的原因或访问的理由。读者随后会比较这些类别,但可靠性规则通常只检查每个估计本身,而不检查它们之间的差异。我们转而询问在抽样设计下哪些类别组可以被区分开。这些类别通常分为若干块,例如与驾驶员相关的原因和与车辆相关的原因。当一个分组的每一对组在同一块内的差异大于设计的最小可检测差异(根据设计自由度下的初级抽样单元计算)时,该分组被称为可区分的。最大可区分分组是最详细的,且其大小可能不同,而合并两个组可能会破坏可区分性。因此,即使将任何一个组一分为二会破坏可区分性,仍可能存在更详细的可区分分组。我们提出了两种方法。一种方法是将类别合并直到分组可区分,然后测试每一个单独的分割。另一种方法列出具有少量类别的块的所有可区分分组。当不同块中的组也必须不同时,可能没有符合条件的分组。对具有12层中24个初级抽样单元的国家机动车碰撞原因调查的应用,给出了其两个类别块的分组最大值。在其13个与车辆相关的关键原因中,超过2700万个分组中,没有超过四个组的分组是可区分的,而其与环境相关的原因最多分为四组。合并一个可区分车辆分组的两个组,在31,773种情况中有5,239种破坏了可区分性。如果不同块的组也必须被区分,则没有分组符合条件。分析人员可以在发布来自自由度较少设计的分类比较之前应用这些方法。
英文摘要
Surveys often estimate the share of cases in each of several categories, such as the causes of an event or the reasons for a visit. Readers then compare the categories, but reliability rules usually check each estimate alone, not the differences between them. We ask instead which groups of categories can be told apart under the sampling design. The categories often fall into blocks, such as driver-related and vehicle-related causes. A grouping is called distinguishable when every pair of its groups in the same block differs by more than the design's minimum detectable difference, computed from the primary sampling units at the design's degrees of freedom. The maximal distinguishable groupings are the most detailed and can differ in size, while merging two groups can destroy distinguishability. So even if splitting any one group in two breaks distinguishability, a more detailed distinguishable grouping may still exist. We propose two methods. One merges categories until the grouping is distinguishable, then tests every single split. The other lists every distinguishable grouping of a block with few categories. When groups in different blocks must also differ, there might be no grouping that qualifies. An application to the National Motor Vehicle Crash Causation Survey, with 24 primary sampling units in 12 strata, gives the grouping maxima for two blocks of its categories. Of more than 27 million groupings of its 13 vehicle-related critical reasons, none with more than four groups is distinguishable, and its environment-related reasons resolve into at most four groups. Merging two groups of a distinguishable vehicle grouping destroys distinguishability in 5,239 of 31,773 cases. If groups from different blocks must also be told apart, no grouping qualifies. Analysts can apply these methods before publishing categorical comparisons from designs with few degrees of freedom.