AI 中文总结
研究针对贝叶斯组序贯设计在设计阶段操作特征评估计算负担大的问题,引入共轭混合半模拟框架,通过创新方法和预计算策略,快速评估校准,应用于ADRENAL试验时大幅提速,使常规校准在计算上可行。
AI 中文摘要
贝叶斯组序贯设计(GSDs)扩展了频率主义GSDs,具有可解释的决策制定和外部证据借用功能,但其在设计阶段操作特征评估的计算负担限制了其应用。传统方法在每次中期观察时使用马尔可夫链蒙特卡罗模拟虚拟试验或近似分析后验更新,使得在商品硬件上联合校准决策阈值和设计框架不切实际。本文引入了一个具有两项创新的半模拟框架。首先,有限共轭混合先验用封闭形式的共轭更新和低维数值积分来计算决策规则尾部概率,取代每次观察(“per-look”)的后验计算。其次,一种预计算策略在所有候选分析时间的并集处通过单次蒙特卡罗遍历缓存per-look后验尾部概率,校准网格中的每个设计通过亚秒级扫描与相同缓存进行评估,无需进一步的模拟成本。该框架支持具有多种疗效和无效标准的后验概率决策规则,在有约束力或无约束力的无效情况下均可使用,并为二元、连续、计数和事件发生时间终点推导封闭形式的per-look更新,这里的基准测试集中在二元终点上。当应用于重新设计ADRENAL试验(最多进行九次分析)时,该框架在蒙特卡罗误差范围内再现了BATSS和adaptr的操作特征,同时运行时,每个GSD比adaptr快约7倍至16倍(在百万次试验校准预算下快数百倍),比BATSS快3700倍至6600倍。这使得常规的贝叶斯GSD校准在计算上可用于确证性试验。
英文摘要
Bayesian group sequential designs (GSDs) extend frequentist GSDs with interpretable decision-making and external evidence borrowing, but their use is limited by the computational burden of design-stage operating-characteristic evaluation. Conventional methods simulate virtual trials with Markov chain Monte Carlo or approximate analytical posterior updates at each interim look, making joint calibration of decision thresholds and design skeletons impractical on commodity hardware. Here we introduce a semi-simulation framework with two innovations. First, finite conjugate-mixture priors replace the posterior computation for each look (``per-look'') with closed-form conjugate updates and low-dimensional numerical integration for decision-rule tail probabilities. Second, a precomputation strategy caches per-look posterior tail probabilities from a single Monte Carlo pass at the union of all candidate analysis times, and each design in the calibration grid is evaluated against the same cache by a sub-second sweep, with no further simulation cost. The framework supports posterior-probability decision rules with multiple efficacy and futility criteria under either binding or non-binding futility, and derives closed-form per-look updates for binary, continuous, count and time-to-event endpoints, with benchmarking here focused on the binary endpoint. When applied to re-design the ADRENAL trial, with up to nine analyses, the framework reproduces the operating characteristics of BATSS and adaptr within Monte Carlo error while running, per GSD, approximately $7\times$ to $16\times$ faster than adaptr at a matched budget (several hundredfold at the million-trial calibration budget) and $3{,}700\times$ to $6{,}600\times$ faster than BATSS. This brings routine Bayesian GSD calibration within computational reach for confirmatory trials.