arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

含聚类结构的向量的随机复杂度

Stochastic complexity of vectors containing cluster structure

Daniel Nicorici, Olli Yli-Harja, Jaakko Astola

arXiv 2609.00084首次发表:更新:

发表机构

Medicel Oy; Institute of Signal Processing, Tampere University of Technology(麦迪赛尔公司; 坦佩雷理工大学信号处理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对含聚类结构向量的随机概率计算问题,提出线性时间的递归公式计算NML模型归一化常数,将原多项式时间复杂度优化为线性,为MDL原理下的聚类相关任务提供了高效理论支撑。

AI 中文摘要

本文研究了使用归一化最大似然(NML)模型计算含聚类结构的编码向量的随机概率(最短码长)的问题,这在基于最小描述长度(MDL)原理的数据聚类中具有重要的理论和实践意义,例如用于估计数据的最佳聚类数和最佳聚类结构。基于NML模型直接计算含聚类结构的向量的最短码长需要相对于向量规模和聚类数的多项式时间。我们通过引入递归公式来有效计算NML模型的归一化常数,证明该问题是可处理的;新公式的时间复杂度为线性,与之前相对于向量规模和聚类数的多项式时间形成对比。

英文摘要

This paper studies the problem of computing the stochastic probability (shortest code length) of the encoded vectors containing cluster structure using Normalized Maximum Likelihood (NML) model. This is of great theoretical and practical importance in data clustering based on Minimum Description Length (MDL) principle, such as for estimating the best number of clusters and best cluster structure for the data. Straightforward computation of the shortest code length of the vector containing cluster structure based on the NML model requires polynomial time with respect to the size of the vector and number of clusters. We show that this is a tractable problem by introducing a recursion formula for the efficient computation of normalizing constant from the NML model. The time complexity of the new formula is linear opposed to previous polynomial time with respect to the size of the vector and number of clusters.

Comments8 pages, 2 figures. Originally published in the Proceedings of the International Workshop on Nonlinear Signal and Image Processing (NSIP 2007), Bucharest, Romania, 10-12 September 2007, pp. 164-169

Journal refProceedings of NSIP 2007 - International Workshop on Nonlinear Signal and Image Processing (2007), pp. 164-169

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑