arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33349stat.ME

面向含异常批次异构流数据的可再生在线期望回归

Renewable Online Expectile Regression for Heterogeneous Streaming Data with Abnormal Batches

  • School of Economics and Management, Beihang University(北京航空航天大学经济管理学院)
  • MOE Key Laboratory of Complex System Analysis and Management Decision, Beihang University(北京航空航天大学复杂系统分析与管理决策教育部重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Wei Cao, Shanshan Wang

AI总结:

针对流数据中异常批次导致的批次异质性,提出基于检测和自适应加权的可再生在线期望回归方法,并引入Huber损失增强稳健性,提升估计精度与鲁棒性。

AI中文摘要:

流数据具有数据量大、到达速度快和分布不断变化的特点,在现代应用中日益普遍。因此,开发高效可靠的估计程序对于实时统计分析至关重要。然而,大多数现有的在线估计方法依赖于批次同质性的假设,而在实际中,由于异常批次、分布偏移或其他形式的批次异质性,这一假设可能被违反。为了解决这一挑战,我们针对异构流数据开发了可再生在线期望回归程序。具体而言,我们提出了两种互补的策略来处理异常批次:(1)一种基于检测的方法,采用基于得分检验统计量的序贯监测机制来识别并移除潜在的异常批次;(2)一种自适应加权方法,为到达的批次分配数据驱动的权重,从而减少异常或漂移批次的影响,同时保留来自可靠观测的信息。这两种策略仅依赖于得分检验统计量,并且可以无缝集成到现有的可再生估计和推断框架中,而无需额外的结构假设。此外,为了增强对重尾误差和异常值的稳健性,我们用Huber损失替代传统的l2损失,并开发了可再生在线期望回归的稳健扩展。大量的模拟研究和临床数据集分析表明,所提出的方法在批次异质性存在的情况下实现了更高的估计精度和稳健性。总体而言,所提出的框架为复杂流数据环境中的可再生期望回归提供了一种灵活且有效的解决方案。

英文摘要:

Streaming data, characterized by high volume, rapid arrival rates, and evolving distributions, have become increasingly prevalent in modern applications. Developing efficient and reliable estimation procedures is therefore essential for real-time statistical analysis. However, most existing online estimation methods rely on the assumption of batch homogeneity, which can be violated in practice due to abnormal batches, distributional shifts, or other forms of batch heterogeneity. To address this challenge, we develop renewable online expectile regression procedures for heterogeneous streaming data. Specifically, we propose two complementary strategies for handling abnormal batches: (1) a detection-based approach that employs a sequential monitoring mechanism based on score test statistics to identify and remove potentially abnormal batches; and (2) an adaptive-weighting approach that assigns data-driven weights to incoming batches, reducing the influence of abnormal or drifting batches while retaining information from reliable observations. Both strategies rely solely on score test statistics and can be seamlessly integrated into existing renewable estimation and inference frameworks without requiring additional structural assumptions. Furthermore, to enhance robustness against heavy-tailed errors and outliers, we replace the conventional l2 loss with the Huber loss and develop a robust extension of renew?able online expectile regression. Extensive simulation studies and analyses of clinical datasets demonstrate that the proposed methods achieve improved estimation accuracy and robustness in the presence of batch heterogeneity. Overall, the proposed framework provides a flexible and effective solution for renewable expectile regression in complex streaming data environments.

↑