超越特征重要性:聚类解释中模式检测方法的对比分析
Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation
AI总结:
本研究对比评估Random Forest替代模型、LIME及主成分分析等聚类模式检测方法,发现其无法稳定识别全部注入模式,凸显现有工具的不足,推动专用方法开发。
AI中文摘要:
解释聚类结果仍是数据分析中的基础挑战,尤其在医疗等需从高维数据提取有意义模式的领域。尽管存在大量可解释性技术,它们主要用于评估特征重要性或提供局部实例级解释,而非识别聚类内的结构化模式。本研究对聚类结果模式检测常用的事后分析方法开展对比评估,为实现可控评估,引入一套系统注入了预定义模式的合成数据集,评估了三种广泛使用的技术:带置换特征重要性的Random Forest替代模型、LIME(Local Interpretable Model-agnostic Explanations)及主成分分析。结果表明,尽管每种方法均能成功恢复相关特征,但 none 始终无法检测所有注入的模式类型。这些发现凸显了现有可解释性工具与模式级聚类解释需求间的关键差距,推动专用模式检测方法的开发。
英文摘要:
Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they are primarily designed to assess feature importance or provide local instance-level explanations rather than to identify structured patterns present within clusters. This work presents a comparative evaluation of commonly used post-hoc analysis methods for pattern detection in clustering results. To enable controlled evaluation, we introduce a suite of synthetic datasets in which predefined patterns are systematically injected. Three widely used techniques are evaluated: a Random Forest surrogate model with permutation feature importance, LIME (Local Interpretable Model-agnostic Explanations), and principal component analysis. Results demonstrate that although each method can successfully recover relevant features, none consistently detects all injected pattern types. These findings high- light a critical gap between existing explainability tools and the requirements of pattern-level cluster interpretation, motivating the development of dedicated pattern detection methodologies.