arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27248cs.LG

重新利用预训练的大语言模型作为高保真连续文本自编码器

Repurposing Pre-trained LLMs as High Fidelity Continuous Text Autoencoders

Arkanath Pathak, Unnat Jain, Alexander C. Berg

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出LLMAE,通过暴露中间固定长度潜瓶颈,将预训练解码器语言模型重用作连续文本自编码器,实现高保真重建并支持下游潜文本扩散模型用于图像描述。

中文摘要 AI 辅助

下一个词元预测使得自回归语言模型具有高度流畅性,但它仅通过序列分解间接地表示全局结构。相比之下,高保真自编码器已成为图像生成中的标准原语,使生成模型能够在连续潜空间上操作;而文本缺乏同样忠实的连续表示。我们提出LLMAE,一种通过在其内部激活中暴露一个中间固定长度的潜瓶颈,将预训练的仅解码器语言模型重新用作连续文本自编码器的方法。LLMAE采用参数高效的270M Gemma 3模型实例化,利用结构化注意力掩码、LoRA适配和KL正则化来学习一个自编码接口,该接口利用了原始大语言模型的生成先验。我们训练LLMAE重建长达1024个词元的文本序列,在该任务上显著改进,实现了近乎完美的重建。此外,我们通过使用学到的LLMAE自编码器训练一个用于详细图像描述的潜文本扩散模型,展示了该表示的下游实用性。通过将文本映射到固定长度的连续潜空间,我们的方法为下游适应提供了有效的基底,同时受益于原始大语言模型的流畅性。

英文摘要

Next-token prediction has enabled highly fluent autoregressive language models, but it represents global structure only indirectly through sequential factorization. In contrast, high-fidelity autoencoders have become a standard primitive in image generation, enabling generative models to operate over continuous latent spaces; text lacks a comparably faithful continuous representation. We propose LLMAE, a method for repurposing a pretrained decoder-only language model as a continuous text autoencoder by exposing the activations of an intermediate layer as a fixed-length latent bottleneck. LLMAE achieves this interface with structured attention masks and LoRA adaptation, leveraging the generative prior of the original LLM. Across two backbones (270M Gemma 3 and 0.5B Qwen2.5), LLMAE reconstructs sequences up to 1024 tokens with high accuracy, reproducing up to 97% of documents verbatim. We demonstrate downstream utility by training a lightweight diffusion model that generates detailed image captions directly in the frozen LLMAE latent space. By mapping text into a fixed-length continuous latent space, our approach provides an effective substrate for downstream adaptation while benefiting from the fluency of the original LLM. Code and models are available at https://github.com/arkanath/LLMAE.

发表机构

  • University of California, Irvine(加州大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

↑