AI 中文总结
研究针对当前多模态语言模型缺乏重建性记忆的问题,提出DoYouRemember三阶段架构,通过VQ-VAE、LoRA微调语言模型和扩散解码器实现,经实验发现语言模型隐藏状态视觉信息少,还解决了共享记忆矩阵训练问题,统一相关发现于信息论框架。
AI 中文摘要
人类记忆是重建性的,而非如实记录。当前多模态语言模型(MLLMs)缺乏此能力:其通过冻结的视觉编码器处理图像,生成一次性文本输出并丢弃内部表示。我们提出了DoYouRemember,这是一种三阶段架构,将重建性记忆引入MLLMs:(1)一个VQ-VAE将图像压缩为离散视觉令牌;(2)一个LoRA微调的语言模型联合处理视觉和文本令牌;(3)一个扩散解码器从语言模型的隐藏状态重建图像。在1000个3D面部皮肤纹理图和99000个未标记面部图像上,我们发现语言模型隐藏状态包含几乎为零的可恢复视觉信息。训练共享记忆矩阵M在反向传播下因梯度抵消而系统失败。我们识别出三个根本原因,并表明局部指数移动平均(EMA)更新可解决所有三个问题。由此产生的M(229K参数,16倍压缩)在未见测试图像上接近VQ上限。扩展到1024个插槽则超过了它。我们在信息论框架下统一了这些发现:记忆是有损压缩,回忆是解压缩,幻觉是有损解压缩的固有属性而非缺陷。
英文摘要
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through a frozen visual encoder, produce a one-shot text output, and discard internal representations. We present DoYouRemember, a three-stage architecture introducing reconstructive memory into MLLMs: (1) a VQ-VAE compresses images into discrete visual tokens, (2) a LoRA-fine-tuned LLM jointly attends to visual and text tokens, and (3) a Diffusion Decoder reconstructs images from the LLM's hidden states. On 1,000 3D facial skin texture maps and 99,000 unlabeled facial images, we find that LLM hidden states contain approximately zero recoverable visual information -- the same Decoder producing clear reconstructions from VQ-VAE tokens (pre-LLM) produces pure noise from LLM hidden states (post-LLM), demonstrating that the LLM understands images but does not remember them. Training a shared memory matrix M under backpropagation systematically fails due to gradient cancellation (O(1/sqrt(N)) attenuation). We identify three root causes and show that local EMA updating resolves all three: each image updates only its top-8 slots out of 64, preserving inter-slot diversity. The resulting M (229K parameters, 16x compressed) approaches the VQ upper bound on unseen test images. Scaling to 1,024 slots surpasses it (LPIPS 0.056 vs. 0.071), as M's continuous representation avoids VQ quantization error. We unify these findings under an information-theoretic framework: memory is lossy compression, recall is decompression, and hallucination is an inherent property of lossy decompression rather than a defect.
Comments43 pages, 8 figures, 3 tables