AI 中文总结
该研究针对机器学习基准中长时聚合标签与短窗口输入配对的性能上限问题,通过特质-状态高斯过程框架推导贝叶斯风险恒等式,分解标签方差、推导有效时间跨度,揭示标签定义学习的相关时间尺度,区分了协议与架构限制。
AI 中文摘要
机器学习基准通常将聚合了长时间范围的标签,与通过一个或几个短窗口观测到的输入配对。因此,它们表现出的性能上限可能是采集协议的上限,而非模型容量的上限。我们研究了当潜在高斯过程同时包含稳定的个体特质和相关的个体内状态时,形式为$Θ_{g,T}=T^{-1}\int_0^T g\{Z(t)\}\\,\mathrm{d}t$的标签。一个精确的、以协议为条件的贝叶斯风险恒等式提供了通用工具。首先,我们将标签方差分解为$O(1)$的特质分量和$O(T^{-1})$的状态分量,解释了为何快照能保留横截面可预测性,却难以追踪个体内变化。其次,我们推导出依赖于任务的有效时间跨度:均值标签依赖于普通相关时间,而占据时间(occupation-time)标签则依赖于一整个高阶相关时间谱。第三,当稳定特质处于阈值时,状态驱动的占据标签方差达到最大;远离该边界时,窗口效率的衰减要慢得多。在等分段预算下,精确风险和蒙特卡洛实验表明,同一时间的重复分段会迅速饱和,而时间上分散的观测会持续提升状态可解释性。特质上限使用普通重测数据中可获得的量;只有状态上限需要短滞后时间校准。这些结果区分了架构限制和协议限制,并表明相关时间尺度由标签定义,而非仅由时长或分段数量定义。
英文摘要
Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows. Their apparent performance ceiling may therefore be an acquisition-protocol ceiling rather than a model-capacity ceiling. We study labels of the form $Θ_{g,T}=T^{-1}\int_0^T g\{Z(t)\}\,\mathrm{d}t$ when the latent Gaussian process contains both a stable individual trait and a correlated within-individual state. An exact protocol-conditioned Bayes-risk identity provides a common tool. First, we decompose label variance into an $O(1)$ trait component and an $O(T^{-1})$ state component, explaining why a snapshot can retain cross-sectional predictability while poorly tracking within-person change. Second, we derive task-dependent effective temporal spans: mean labels depend on the ordinary correlation time, whereas occupation-time labels depend on an entire spectrum of higher-order correlation times. Third, state-driven occupation-label variance is maximal when the stable trait lies at the threshold; window efficiency decays much more slowly away from that boundary. Under an equal segment budget, exact risks and Monte Carlo experiments show that repeated segments at one time rapidly saturate, whereas temporally dispersed observations continue to increase state explainability. The trait ceiling uses quantities available from ordinary test-retest data; only the state ceiling requires short-lag temporal calibration. The results distinguish architectural limits from protocol limits and show that the label, rather than duration or segment count alone, defines the relevant timescale.