发表机构
Harbin Engineering University, China.(哈尔滨工程大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有最优传输方法在短文本聚类中忽视语义一致性的问题,提出新框架,设计实例级注意力机制并融入OT公式,获取兼顾语义一致性和全局结构信息的伪标签,提升短文本聚类效果,优于现有方法。
AI 中文摘要
基于最优传输(OT)的伪标签已成为增强短文本聚类的有效机制。现有OT方法在建模样本间语义一致性方面不足,可能给语义相似样本分配不同伪标签,导致模型产生劣质聚类。本文提出一种新颖的短文本聚类框架,弥补现有OT方法对语义一致性的忽视,生成可靠伪标签促进聚类。具体而言,该方法首先设计实例级注意力机制捕捉样本间语义关系,将其融入OT公式赋予传输过程邻域语义感知。通过求解提出的OT公式,获得同时考虑样本间语义一致性和样本与聚类全局结构信息的可靠伪标签,用作监督信号引导模型实现准确聚类。大量实验表明该方法优于现有方法。代码可在指定链接获取。
英文摘要
Pseudo-labeling based on Optimal Transport (OT) has become an effective mechanism for enhancing short text clustering. Existing OT methods are short in modeling semantic consistencies between samples, which may assign different pseudo-labels to semantically similar samples. These erroneous pseudo-labels can cause the model to produce inferior clusters. This paper proposes a novel short text clustering framework, which remedies the neglect of semantic consistency in existing OT methods, generating reliable pseudo-labels to facilitate clustering. Specifically, the proposed approach first designs an instance-level attention mechanism to capture semantic relationships between samples, which are then integrated into the OT formulation to endow the transport process with neighborhood semantic awareness. By solving the proposed OT formulation, reliable pseudo-labels are obtained that simultaneously account for sample-to-sample semantic consistency and sample-to-cluster global structure information. These pseudo-labels are then used as supervisory signals to guide the model to achieve accurate clustering. Extensive experiments demonstrate that the proposed approach outperforms state-of-the-art methods. The code is available at: \href{https://github.com/YZH0905/CAOT-STC}{https://github.com/YZH0905/CAOT-STC}.