arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21538stat.ME

模糊熵三重k均值

Fuzzy entropic triple k-means

  • Sapienza University(罗马第一大学)

机构由 AI 辅助整理,请以论文原文为准。

Mariaelena Bottazzi Schenone, Roberto Rocci, Maurizio Vichi

中文总结 AI 辅助

提出模糊熵三重k均值(FE3KM),通过熵正则化线性嵌入隶属度,实现三维数据的同时聚类,并支持偏差分解与解释,优于传统模糊聚类。

中文摘要 AI 辅助

本文提出模糊熵三重k均值(FE3KM),一种用于三维数据数组的熵正则化模糊划分方法,可同时聚类对象、变量和场合。与通过非线性指数m控制模糊度的标准模糊聚类不同,FE3KM在最小二乘(LS)目标中线性嵌入隶属度,使用行随机隶属度矩阵和熵正则化来表示聚类分配中的不确定性。这种线性结构使得总偏差能够精确分解为簇内和簇间成分,并可进一步归因于每个模式(对象、变量、场合)以及每个模式内的各个簇。这是指数m模糊模型所不具备的解释性属性。我们推导了FE3KM的交替最小二乘算法,包括用于大型数组的加速变体,并正式将该方法与串联聚类程序以及Tucker3/三模式划分模型联系起来。一项模拟研究评估了恢复准确性、对噪声和模糊度错误设定的鲁棒性,以及与串联和三路聚类基线的性能对比。对基准三维电视收视率数据集的应用展示了FE3KM如何恢复可解释的对象-变量-场合结构,并量化每个簇和模式对总体数据变异性的贡献。

英文摘要

This paper proposes fuzzy entropic triple k-means (FE3KM), an entropy-regularized fuzzy partitioning method for three-way data arrays that simultaneously clusters objects, variables, and occasions. Unlike standard fuzzy clustering, which controls fuzziness through a nonlinear exponent m, FE3KM embeds memberships linearly within a Least-Squares (LS) objective, using row-stochastic membership matrices and entropy regularization to represent uncertainty in cluster assignments. This linear structure is what enables an exact decomposition of total deviance into within- and between-cluster components, further attributable to each mode (objects, variables, occasions) and to individual clusters within each mode. This is an interpretive property unavailable to exponent-m fuzzy models. We derive Alternating Least-Squares algorithms for FE3KM, including an accelerated variant for large arrays, and formally relate the method to tandem clustering procedures and to Tucker3/three-mode partitioning models. A simulation study assesses recovery accuracy, robustness to noise and fuzziness mis-specification, and performance against tandem and three-way clustering baselines. An application to a benchmark three-way TV ratings dataset demonstrates how FE3KM recovers interpretable object-variable-occasion structures and quantifies the contribution of each cluster and mode to overall data variability.

补充信息

↑