arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于带相关误差的狄利克雷过程混合模型的函数型数据聚类变分推断

Variational Inference for Functional Data Clustering via Dirichlet Process Mixtures with Correlated Errors

Chengqian Xian

arXiv 2609.04853首次发表:更新:

AI 中文总结

该研究提出一种基于贝叶斯模型的函数型数据聚类方法,采用B样条、Ornstein--Uhlenbeck结构与截断狄利克雷过程混合模型,通过变分EM算法实现高效后验近似,模拟与实际应用验证了其良好性能。

AI 中文摘要

我们提出一种基于贝叶斯模型的方法,用于聚类具有未知簇数量和时间相关观测的函数型数据。簇特异性均值函数采用B样条基展开表示,而曲线内的相关性通过Ornstein--Uhlenbeck协方差结构建模。我们使用截断狄利克雷过程混合模型推断有效簇数量,并开发了一种变分EM算法用于高效的后验近似。模拟研究表明,所提方法在均值函数设定正确和设定错误的情况下均表现良好,且与几种现有函数型聚类方法相比具有竞争力。与MCMC的比较表明,变分近似能产生一致的聚类和参数估计,但计算成本显著更低。将其应用于加拿大日温度曲线进一步证明了该方法的实用性,其可在考虑时间相关性的同时识别可解释的函数型簇。

英文摘要

We propose a Bayesian model-based approach for clustering functional data with an unknown number of clusters and within-curve correlated observations. Cluster-specific mean functions are represented using B-spline basis expansions, while within-curve dependence is modeled through an Ornstein--Uhlenbeck covariance structure. A truncated Dirichlet process mixture is used to infer the effective number of clusters, and a variational EM algorithm is developed for efficient posterior approximation. Simulation studies show that the proposed method performs well under both correctly specified and misspecified mean-function settings and achieves higher average values of the reported clustering metrics than the competing methods in the simulation settings considered. Comparisons with MCMC indicate that the variational approximation produces clustering results and parameter estimates that are in close agreement with those obtained by MCMC, while requiring substantially lower computational cost. An application to Canadian daily temperature curves further demonstrates the practical usefulness of the method in identifying interpretable functional clusters while accounting for within-curve dependence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑