AI 中文总结
该研究提出DPitG方法,结合精度目标与明确裁决,将PitG的不确定率从62%降至2%且零假阳性,符合预注册标准,适用于需可靠裁决的场景。
AI 中文摘要
序贯假设检验相较于固定样本设计具有灵活性,但结合决策准则的停止规则可能因提前观测而产生确认偏差。HDI+ROPE算法便是一例:一旦后验最高密度区间(HDI)完全落在原假设周围的实际等价区间(ROPE)之外(拒绝原假设)或完全落在其中,便会停止,这使得能明确接受原假设,而这是零假设显著性检验无法做到的。然而,这种结合意味着宽HDI可能在早期非代表性样本上满足准则,导致系统性假阳性成本。完全解耦的方法,如「精度即目标」(PitG),仅在达到目标HDI宽度ω时才停止,从而消除这种偏差;但通过将停止规则与决策准则分离,PitG常产生不确定结果,尤其是当原假设为真时。我们提出「决断精度即目标」(DPitG),要求精度目标和明确裁决同时满足。在公平硬币模拟中(ω=0.08,ROPE=0.5±0.05),DPitG将PitG的不确定率从62%降至2%,中位样本量仅增加5%,且零假阳性;HDI+ROPE仅以假阳性为代价达到相当的决断性。在测试的ω范围内,DPitG的决断性保持在97%以上,而PitG则波动较大。我们提供闭式规划公式N∝V/ω²(V为观测方差)、在线交互式计算器及开源代码。该框架在单组二元数据上验证,可轻松扩展至连续结果和两组比较。DPitG完全预先指定的停止规则符合预注册标准,是需要可靠裁决时的首选方法。
英文摘要
Sequential hypothesis testing offers flexibility over fixed-sample designs, but stopping rules coupled to decision criteria risk confirmation bias through early peeking. The HDI+ROPE algorithm exemplifies this: it stops as soon as the posterior Highest Density Interval (HDI) falls entirely outside a Region of Practical Equivalence (ROPE) around the null (rejecting it) or entirely inside, which enables positive acceptance of the null, something Null Hypothesis Significance Testing cannot do. However, this coupling means a wide HDI can satisfy the criterion on early, unrepresentative samples, incurring a systematic false-positive cost. Fully decoupled methods, such as ``Precision is the Goal'' (PitG), eliminate this bias by halting only once a target HDI width $ω$ is reached; yet, by divorcing the stopping rule from the decision criterion, PitG frequently yields inconclusive outcomes, particularly when the null hypothesis is true. We propose ``Decisive Precision is the Goal'' (DPitG), which requires both the precision target and a conclusive verdict to be satisfied simultaneously. In fair coin simulations ($ω=0.08$, ROPE$=0.5\pm0.05$), DPitG reduces the PitG inconclusive rate from 62% to 2% at a median cost of only 5% more samples, with zero false positives; HDI+ROPE achieves comparable conclusiveness only at the cost of false positives. Across the $ω$ range tested, DPitG conclusiveness remains above 97% while PitG's varies widely. We provide a closed-form planning formula $N\propto V/ω^2$ (where $V$ is the observation variance), an online interactive calculator, and open-source code. Demonstrated on single-group binary data, the framework extends readily to continuous outcomes and two-group comparisons. DPitG's fully pre-specified stopping rule is compatible with pre-registration standards and is the method of choice whenever a reliable verdict is required.
CommentsMain section 17 pages and 9 figures, Seven Abstracts in 7 pages and 1 figure