TextEconomizer:利用去噪变换器和熵编码增强有损文本压缩
TextEconomizer: Enhancing Lossy Text Compression with Denoising Transformers and Entropy Coding
- United International University(联合国际大学)
- BRAC University(BRAC大学)
- Southeast University(东南大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出TextEconomizer编码器-解码器框架,结合去噪变换器和熵编码,实现50%-80%的压缩率,参数减少153倍,在BLEU等指标上保持近完美文本质量。
AI中文摘要:
有损文本压缩在保留核心含义的同时减少数据大小,适用于摘要、自动分析和数字存档。尽管基于变换器的模型在语言建模中占主导地位,但将上下文向量和熵编码集成到序列到序列(Seq2Seq)生成中仍未充分探索。一个关键挑战在于从编码器输出中识别信息最丰富的上下文向量,并引入熵编码以提高存储效率,同时即使在噪声文本下也能保持高质量输出。我们提出了TextEconomizer,一种与变换器神经网络配对的编码器-解码器框架,无需数据集维度的先验知识即可将可变大小输入减少50%至80%。我们的模型通过熵编码实现了有竞争力的压缩比,同时通过BLEU、ROUGE、METEOR和语义相似度评分评估,提供了近乎完美的文本质量。TextEconomizer的参数数量比同类模型少约153倍,实现了5.39倍的压缩比,且不牺牲语义质量。我们还评估了一个基于LSTM的自编码器,实现了最先进的67倍压缩比,参数减少196倍;以及LLaMAFormer,一种改进的变换器,参数比ICAE少263倍,同时保持有竞争力的文本质量。TextEconomizer在平衡内存效率和高保真输出方面显著超越了现有的基于变换器的模型,标志着有损压缩在最优空间利用方面的突破。
英文摘要:
Lossy text compression reduces data size while preserving core meaning, making it well-suited for summarization, automated analysis, and digital archives. Despite the dominance of transformer-based models in language modeling, integrating context vectors and entropy coding into Sequence-to-Sequence (Seq2Seq) generation remains underexplored. A key challenge lies in identifying the most informative context vectors from encoder output and incorporating entropy coding to enhance storage efficiency while maintaining high-quality outputs, even under noisy text. We introduce TextEconomizer, an encoder-decoder framework paired with a transformer neural network that reduces variable-sized inputs by 50% to 80% without prior knowledge of dataset dimensions. Our model achieves competitive compression ratios via entropy coding while delivering near-perfect text quality, assessed by BLEU, ROUGE, METEOR, and semantic similarity scores. TextEconomizer operates with approximately 153x fewer parameters than comparable models, achieving a 5.39x compression ratio without sacrificing semantic quality. We also evaluate an LSTM-based autoencoder achieving a state-of-the-art 67x compression ratio with 196x fewer parameters, and LLaMAFormer, a modified transformer with 263x fewer parameters than ICAE while maintaining competitive text quality. TextEconomizer significantly surpasses existing transformer-based models in balancing memory efficiency and high-fidelity outputs, marking a breakthrough in lossy compression with optimal space utilization.