动态语义压缩:面向大语言模型高效潜在空间推理
Dynamic Semantic Compression for Efficient Latent-Space Inference in Large Language Models
浏览论文内容
中文总结 AI 辅助
提出DSEI框架,通过动态语义自编码器将LLM推理从词元级转为潜在空间片段级,在Wanjuan数据集上困惑度降低48%,推理速度提升2.5倍,内存开销降低90%。
中文摘要 AI 辅助
大语言模型(LLMs)主要在词元级别进行推理,导致大量内存开销和计算效率受损。本文提出一种动态语义提取与推理(DSEI)框架,通过两阶段训练策略在潜在空间内实现片段级推理。首先,我们通过自监督学习构建动态语义自编码器(DSAE)。DSAE通过自适应语义加权和门控融合动态提取片段级语义,并将其压缩为紧凑的潜在表示。随后,我们将DSAE集成到LLM架构中,训练模型在稠密潜在空间上进行推理。DSEI大幅缩短了输入和生成序列的长度,显著提升了推理效率。在Wanjuan数据集上进行的大量实验表明,与静态句子级潜在推理基线相比,DSEI将困惑度降低了48%。此外,与采用词元级推理的标准LLM相比,DSEI将推理速度提升了2.5倍,并将内存开销降低了90%。
英文摘要
Large Language Models (LLMs) primarily perform inference at the token level, resulting in substantial memory overhead and compromised computational efficiency. In this paper, we propose a Dynamic Semantic Extraction and Inference (DSEI) framework, which achieves segment-level inference within the latent space through a two-stage training strategy. First, we construct a Dynamic Semantic Autoencoder (DSAE) via self-supervised learning. DSAE dynamically extracts segment-level semantics and compresses them into compact latent representations via adaptive semantic weighting and gated fusion. Subsequently, we integrate the DSAE into the LLM architecture and train the model to infer over dense latent space. DSEI substantially reduces both input and generation sequences and significantly enhances inference efficiency. Extensive experiments conducted on the Wanjuan dataset demonstrate that DSEI reduces perplexity by 48% compared to static sentence-level latent inference baseline. Furthermore, compared to standard LLMs using token-level inference, DSEI accelerates inference speed by 2.5$\times$ and reduces memory overhead by 90%.
发表机构
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。