arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于同质性-简约性权衡的外部聚类验证

External Clustering Validation by the Homogeneity-Parsimony Trade-off

Andreas Tiffeau-Mayer

arXiv 2607.20799首次发表:更新:

发表机构

Division of Infection and Immunity & Institute for the Physics of Living Systems, University College London(伦敦大学学院感染与免疫系及生命系统物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究如何通过同质性-简约性权衡进行外部聚类验证,基于信息瓶颈原理构建分数,推导相关对应物统一评估标准,展示该框架在特征选择和算法比较中的效用,能阐明聚类操作点并识别帕累托最优解。

AI 中文摘要

标量指标常用于根据已知类别评估聚类,但它们可能掩盖一个基本权衡:聚类应能提供有关类别标签的信息,同时避免不必要的碎片化。本文描述了量化此权衡的聚类同质性和简约性的归一化分数。这些分数基于信息瓶颈原理构建,并进行了修改以不奖励有损压缩。通过示例和数学证明表明,与相关提议相比,我们对这些分数的定义在聚类细化下具有单调变化的直观属性。扩展信息理论框架超出香农熵,我们还推导了同质性和简约性分数的集匹配和基于对的对应物。这些统一了常用评估标准,并表明在基于对的设置中,同质性-简约性权衡恢复了二元分类器的接收者操作特征。我们展示了该框架在特征选择和算法比较中的效用,说明了联合考虑分数如何能阐明聚类操作点并识别帕累托最优解。

英文摘要

Scalar metrics are often used to evaluate clusterings against known classes, but they can obscure a fundamental trade-off: clusterings should be informative about class labels while avoiding unnecessary fragmentation. Here we describe normalized scores of cluster homogeneity and parsimony that quantify this trade-off. These scores build on the information bottleneck principle, modified to not reward lossy compression. We show by example and mathematical proof that our definitions of these scores have the intuitive property of varying monotonically under cluster refinement in contrast to related proposals. Extending the information-theoretic framework beyond Shannon entropies, we furthermore derive set-matching and pair-based counterparts of the homogeneity and parsimony scores. These unify commonly used evaluation criteria and show that, in the pair-based setting, the homogeneity-parsimony trade-off recovers the receiver operating characteristic of binary classifiers. We demonstrate the framework's utility for feature selection and algorithm comparison, illustrating how considering scores jointly can clarify clustering operating points and identify Pareto-optimal solutions.

Comments12 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑