arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

依赖序列中罕见事件 onset 预测的最优分层分配

Optimal Stratified Allocation for Rare-Event Onset Forecasting in Dependent Sequences

Jaskaran Singh

arXiv 2609.04420首次发表:更新:

发表机构

Indraprastha Institute of Information Technology Delhi(德里印度普拉斯塔信息技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对依赖序列中罕见事件 onset 预测问题,推导了最优分层分配方案,经实证验证其在价格制度预测中具有良好排序性能,但跨 horizon 的π依赖预测未成立。

AI 中文摘要

设一个包含n个带标签示例的有限总体,其具有类加权损失,其中π*n为罕见正类,权重为N0/N1。我们研究在将K个样本分配给两个层(K0和K1)的设计下,从子样本K << n中估计总风险的问题。我们推导了在类条件无放回抽样下加权风险估计量的精确有限总体方差,并求解了最优分配。类乘数通过不平衡比率放大正层的离散度,导致该比率从最优分配中抵消,使得相等分配而非比例分配成为自然默认选择。简单随机抽样受显式层间项支配;精确偏差恒等式表明,集群代表性选择无通用无偏性保证;Serfling界将分配结果转移到有限候选集上的选择误差。在所实现的截断下,实现的分配比率为γ=min(2π/f,1),与n无关,产生无参数效率预测A(π,f)=γ/[π(1−π)(1+γ)²]。单独的可测性结果限定了带标签示例的记录占用,确定了所需的训练-测试分离,并控制了绝对正则性下对块独立性的偏离。我们在预测统计上爆炸性价格制度的 onset 时测试了这些预测,该制度由Phillips-Shi-Yu程序事后确定,使用2004-2011年的350只美国股票,正行占比不足1%,并设置了五个清除的前向块。四个设计的预测排序成立,且在10天 horizon 下,五个设计点的排序与A的预测完全一致(Spearman秩相关系数ρ=1,精确p=0.0167)。跨 horizon 对π的预测依赖不成立;我们确定了基于设计的论证之外的渠道。

英文摘要

Let a finite population of n labelled examples carry a class-weighted loss, with pi*n in a rare positive class weighted by N0/N1. We study estimation of total risk from a subsample K << n under designs allocating K0 and K1 draws to the two strata. We derive the exact finite-population variance of the weighted risk estimator under class-conditional sampling without replacement and solve for the optimal allocation. The class multiplier inflates positive-stratum dispersion by the imbalance ratio, causing that ratio to cancel from the optimal allocation and making equal, rather than proportional, allocation the natural default. Simple random sampling is dominated by an explicit between-stratum term; an exact bias identity shows that cluster-representative selection has no general unbiasedness guarantee; and a Serfling bound transfers the allocation result to selection error over a finite candidate set. Under the implemented truncation, the realised allocation ratio is gamma=min(2*pi/f,1), independent of n, yielding the parameter-free efficiency prediction A(pi,f)=gamma/[pi(1-pi)(1+gamma)^2]. A separate measurability result bounds the record occupied by a labelled example, determining the required train-test separation and controlling departure from block independence under absolute regularity. We test these predictions on forecasting the onset of statistically explosive price regimes, dated ex post by the Phillips-Shi-Yu procedure, using 350 U.S. equities from 2004-2011 with under 1% positive rows and five purged forward blocks. The predicted ordering of the four designs holds, and at the 10-day horizon the five design points are ordered exactly as predicted by A (Spearman rho=1, exact p=0.0167). The predicted dependence on pi across horizons does not hold; we identify the channels lying outside the design-based argument.

Comments24 pages, 11 tables, no figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑