GRACE:用于可扩展混合数据聚类的大语言模型(LLM)基础语义度量空间
GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
- Guangdong University of Technology(广东工业大学)
- Tsinghua University(清华大学)
- Peking University(北京大学)
- Shenzhen University(深圳大学)
- Xiamen University(厦门大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对混合数据聚类中语义丰富度与可扩展性的矛盾,提出GRACE框架,通过多视角LLM查询实现一次性语义嵌入,兼顾可扩展性与聚类性能,优于11种对比方法。
AI中文摘要:
聚类混合表格数据需要统一的度量空间,以弥合连续数值测量与离散类别符号之间的固有异质性。传统算法完全依赖数据集内部统计来估计类别关系,这使得学习到的度量局限于经验共现,忽略了概念上明显但统计上未观测到的关联。尽管大语言模型(LLM)提供外部世界知识,但将其以文本为中心的推理应用于高度抽象的表格概念存在重大挑战。弥合这种模态差距以构建语义完整的度量,通常需要将LLM嵌入迭代度量学习循环以动态优化跨模态表示,这会产生难以处理的计算开销,迫使在语义丰富度和可扩展性之间做出妥协。因此,我们提出GRACE,一个用于可扩展混合数据聚类的LLM基础框架。GRACE通过多视角LLM查询策略将语义获取转移到属性-值层面,将异质值映射为知识驱动的描述。关键是,这种一次性基础提取了通用语义表示,将异质属性嵌入统一空间,将昂贵的LLM调用与迭代优化解耦。此外,GRACE将这些外部语义与数据集内部统计证据交叉验证,以确保与特定数据集的聚类结构对齐。最终,GRACE达到了传统统计驱动基线的可扩展性,同时在11种竞争方法上实现了更优的聚类准确性和概念可解释性。源代码可在指定URL获取。
英文摘要:
Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurements and discrete categorical symbols. Traditionally, algorithms rely entirely on dataset-internal statistics to estimate categorical relationships, which confines the learned metric to empirical co-occurrences and ignores conceptually obvious yet statistically unobserved affinities. Although LLMs offer external world knowledge, applying their text-centric reasoning to highly abstract tabular concepts presents significant challenges. Bridging this modality gap to construct a semantically complete metric typically requires embedding LLMs into iterative metric learning loops to dynamically optimize cross-modality representations. This incurs intractable computational overhead, forcing a compromise between semantic enrichment and scalability. Therefore, we propose GRACE, an LLM-grounded framework for scalable mixed-data clustering. GRACE shifts semantic acquisition to the attribute-value level via a multi-perspective LLM querying strategy, mapping heterogeneous values into knowledge-informed descriptions. Crucially, this one-shot grounding extracts general-purpose semantic representations that embed heterogeneous attributes into a unified space, decoupling expensive LLM invocation from iterative optimization. Furthermore, GRACE cross-validates these external semantics against dataset-internal statistical evidence to ensure alignment with the dataset-specific cluster structure. Ultimately, GRACE matches the scalability of conventional statistics-driven baselines while achieving superior clustering accuracy and conceptual interpretability over 11 competing methods. The source code is available at https://github.com/develop-yang/GRACE-GRACE-A