有损压缩文本自编码器
Lossy Compressive Text Autoencoders
浏览论文内容
中文总结 AI 辅助
本研究提出带残差低维离散瓶颈的自编码器,在网页文本上实现每字节2.24比特的有损压缩,与无损压缩性能相当,且重构及下游任务表现良好。
中文摘要 AI 辅助
本研究探索数据压缩与表示学习交叉领域中文本的压缩潜在表示学习。我们提出一种自编码器架构,沿时间轴对隐藏表示执行残差下采样与上采样,带有残差低维离散瓶颈。我们针对不同量化方法、训练目标和数据集分析该方法,在不同压缩级别下,从表面级(BLEU)和语义级(基于大语言模型的评判器)评估原始文本与重构文本的相似度,还在下游问答和语义文本相似度基准上评估模型。在网页文本数据上,我们的方法生成的压缩表示每字节达2.24比特,与无损文本压缩算法表现相当,同时具备良好的重构和下游任务性能。
英文摘要
Our work explores learning a compressed latent representation of text, at the intersection of data compression and representation learning. We propose an autoencoder architecture that performs residual downscaling and upscaling of hidden representations along the time axis, with a residual low-dimension discrete bottleneck. We analyze our approach for different quantization methods, training objectives, and datasets. For different levels of compression, we evaluate the similarity between the original and reconstructed text both at the surface-level (BLEU) and at the semantic-level (LLM-based judge). Additionally, we evaluate our models on downstream question-answering and semantic text similarity benchmarks. Our approach results in compressed representations which are on par with lossless text compression algorithms at 2.24 bits per byte on web text data, while having good reconstruction and downstream task performance.
发表机构
- Apple(苹果公司)
- EPFL(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。