arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向降采样横截面市场聚合的近似基于设计的区间:带宽受限金融数据管道的随机化设计

Approximate Design-Based Intervals for Downsampled Cross-Sectional Market Aggregates: A Randomized Design for Bandwidth-Constrained Financial Data Pipelines

Minmin Zeng

arXiv 2610.08801首次发表:更新:

发表机构

AlphaGBM(AlphaGBM)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对金融数据降采样中确定性Top-k选择导致聚合误差不可审计的问题,提出基于未平仓合约量的概率抽样与Hajek估计量,在美港期权数据上显著降低误差并实现近似设计区间。

AI 中文摘要

金融机构通常对横截面期权面板进行降采样,以满足带宽和成本约束。行业默认做法——按未平仓合约量进行确定性Top-$k$选择——无法估计其引入的聚合误差,导致下游风险指标不可审计。我们精确刻画了识别问题。对于确定性规则,若其选择集仅依赖于辅助变量而不依赖于结果(目标盲规则,按未平仓合约量Top-$k$即为一例),则仅凭保留的结果无法识别目标,除非施加界限或建模假设;在结果有界的情况下,目标可部分识别,我们以闭式形式刻画了尖锐的最坏情况区间。我们建议用概率抽样替代Top-$k$,其包含概率与未平仓合约量成比例,并配以Hajek比率估计量和线性化的基于设计的方差估计量,从而为每个聚合量生成近似基于设计的区间。在80个美国交易时段(每天211个标的)和83个香港期权时段(每天131个标的)上的实验表明,对于未平仓合约量加权的隐含波动率,随机化设计将平均绝对误差降低了35%(在10%保留率下,MAE为5.31对比8.15;配对$t=6.68$,Cohen's $d=0.75$,移动块自助法95%置信区间$[1.82, 3.96]$),而Top-$k$可用的尖锐最坏情况区间在oracle界限下比抽样区间宽2.7--6.4倍,在可行界限下宽9--26倍。经验覆盖率在名义95%下为91.9%。我们进一步证明,包含概率必须与目标的影响函数匹配,从而为生产管道提出设计感知框架。

英文摘要

Financial institutions routinely downsample cross-sectional options panels to meet bandwidth and cost constraints. The industry default---deterministic Top-$k$ selection by open interest---provides no estimate of the aggregation error it introduces, rendering downstream risk metrics unauditable. We make the identification problem precise. For deterministic rules whose selected set depends only on auxiliary variables and not on the outcome (target-blind rules, of which Top-$k$-by-open-interest is one), the target is not identified from the retained outcomes alone unless bounds or modelling assumptions are imposed; under outcome bounds it is partially identified, with a sharp worst-case interval we characterise in closed form. We propose replacing Top-$k$ with probability sampling whose inclusion probabilities are proportional to open interest, paired with the Hajek ratio estimator and a linearised design-based variance estimator, yielding approximate design-based intervals around every aggregate. Experiments on 80 US trading sessions (211 tickers/day) and 83 sessions of Hong Kong options (131 tickers/day) show that for OI-weighted implied volatility the randomized design reduces mean absolute error by 35% (MAE 5.31 vs. 8.15 at 10% retention; paired $t=6.68$, Cohen's $d=0.75$, moving-block-bootstrap 95% CI $[1.82, 3.96]$), and the sharp worst-case interval available to Top-$k$ is 2.7--6.4$\times$ wider than the sampling interval under oracle bounds, and 9--26$\times$ wider under feasible bounds. Empirical coverage is 91.9% against a nominal 95%. We further document that inclusion probabilities must match the target's influence function, motivating a design-aware framework for production pipelines.

Comments44 pages, 2 figures, 28 tables. Supplementary material included

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑