发表机构
ABV-Indian Institute of Information Technology and Management(ABV-印度信息技术与管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ICLFS方法,将无监督特征选择转化为特征层面的表示学习问题,在12个基准数据集的实验中,其在10个数据集上的聚类准确率优于各类基线方法。
AI 中文摘要
无监督特征选择(UFS)旨在在不访问类标签的情况下找到一组紧凑的信息特征,这使得特征效用难以定义。因此,现有的UFS方法依赖于间接结构准则,如相似性保持、局部性、稀疏性、聚类几何或重建质量。本文转而通过表示一致性研究UFS,并提出用于无监督特征选择的倒置对比学习(ICLFS),这是一种逐特征的对比框架,将UFS重新表述为针对特征而非样本的表示学习问题。ICLFS首先倒置数据矩阵,使每个特征由其样本轮廓向量表示,随后构建多个掩码正视图和一个打乱的负视图,并在基于InfoNCE的目标下学习在这些结构化扰动下保持一致的投影空间表示。受基于余弦和InfoNCE的训练会影响嵌入范数这一近期发现的启发,我们使用投影空间嵌入幅度作为对特征进行排序的显著性信号。随后,通过拉普拉斯门控排序修正对所得的基于范数的排序进行优化,该修正会抑制局部冗余候选,同时保留显著候选。在12个基准数据集上进行的大量实验表明,在标准的基于聚类的UFS评估协议下,与经典和神经基线方法相比,ICLFS在10个数据集上实现了最佳聚类准确率,在另外两个数据集上也具有竞争力。这些结果表明,逐特征的对比表示一致性为基于邻域、聚类和重建的UFS公式提供了一种强大且有效的替代方案。
英文摘要
Unsupervised feature selection seeks a compact subset of informative features without access to class labels, making feature utility difficult to define. Existing UFS methods therefore rely on indirect structural criteria, such as similarity preservation, locality, sparsity, cluster geometry, or reconstruction quality. In this paper, we instead study UFS through representation consistency and propose Inverted Contrastive Learning for Unsupervised Feature Selection (ICLFS), a feature-wise contrastive framework that reformulates UFS as a representation learning problem over features rather than samples. ICLFS first inverts the data matrix so that each feature is represented by its sample-profile vector, then constructs multiple masked positive views together with a shuffled negative view, and learns projector-space representations that remain consistent across these structured perturbations under an InfoNCE-based objective. Motivated by recent findings that cosine-based and InfoNCE-based training affect embedding norms, we use projector-space embedding magnitude as the saliency signal for ranking features. The resulting norm-based ranking is subsequently refined through Laplacian-Gated Ranking Correction, which suppresses locally redundant candidates while preserving salient ones. Extensive experiments on 12 benchmark datasets show that ICLFS achieves the best clustering accuracy on 10 datasets against both classical and neural baselines under the standard clustering-based UFS evaluation protocol, while remaining competitive on the other two. These results show that feature-wise contrastive representation consistency provides a strong and effective alternative to neighborhood, cluster, and reconstruction-based UFS formulations.