arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CDAE:用对比去噪增强预训练语言模型的扰动鲁棒性

CDAE: Enhancing Perturbation Robustness in Pretrained Language Models with Contrastive Denoising

Sina Heydari, Amirreza Abbasi, Mohsen Hooshmand, Majid Ramezani

arXiv 2607.28236首次发表:更新:

发表机构

Institude for Advanced Studies in Basic Sciences (IASBS)(基础科学高级研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对预训练语言模型词嵌入对保留语义的文本扰动敏感的问题,提出CDAE模型,经多强度扰动策略对比实验,证实其能提升扰动下的嵌入相似度与表示稳定性。

AI 中文摘要

预训练语言模型大幅提升了句子表示学习能力,但其词嵌入对同义词替换、掩码、词丢弃等保留语义的文本扰动仍较为敏感。本研究提出轻量型对比去噪自编码器(CDAE),通过联合优化对比与重构目标来改进预训练BERT的词嵌入,以学习扰动不变的表示。我们采用多种强度不同的扰动策略评估该框架,将其与原始BERT嵌入、SimCSE进行对比。实验结果显示,CDAE在扰动下始终保持更高的嵌入相似度,且随着扰动强度提升,改进效果愈发显著;该框架有效增强了表示稳定性,同时保留语义信息,表明扰动不变学习是提升句子嵌入的有前景方向。源代码公开于指定网址。

英文摘要

Pre-trained language models have significantly improved sentence representation learning, yet their embedding remain sensitive to semantic preserving textual perturbations such as synonym substitution, masking and word dropout. This work proposes a lightweight Contrastive Denoising Autoencoder (CDAE) that refines pre-trained BERT embedding by jointly optimizing contrastive and reconstruction objective to learn perturbation-invariant representation. We evaluate the proposed framework using multiple perturbation strategies with varying strengths and compare it against the original BERT embeddings and SimCSE. Experimental results show that CDAE consistently preserves higher embedding similarity under perturbations, with the improvements becoming more pronounced as framework effectively enhances representation stability while preserving semantic information, highlighting perturbation-invariant learning as a promising direction for improving sentence embeddings. The source code is publicly available at: https://github.com/ComputationIASBS/CDAE

CommentsSubmitted to 16th International Conference on Computer and Knowledge Engineering (ICCKE 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑