AI 中文总结
本研究实证发现训练后模型保持块级权重结构,据此提出LinkerLLM加载器,通过共享块节省GPU内存,支持多模型共驻。
AI 中文摘要
现代LLM作为训练后变体(基础、指令、聊天、代码)的家族部署,这些变体源自共享的预训练权重。我们对训练后如何改变权重空间几何进行了实证研究,覆盖了四个架构家族(Qwen2.5、Llama-3.1/3.2、Mistral、Gemma-2)中的八种配置。我们识别出一个粒度差距:训练后修改了每个张量(291-339个张量中零个保持字节相同,因此基于哈希的去重实现0%的节省),但保留了块级结构(平均余弦相似度超过0.99,相对Frobenius距离保持在0.13以下)。因此,训练后作为一种结构化扰动,移动每个参数同时保持块级几何不变。该属性并非普遍存在:独立训练的专业化模型(例如Qwen2.5-Coder)与通用基础模型的余弦相似度约为0.64,表明权重空间中存在不连通区域。扰动幅度随模型规模、架构家族和训练后配方系统性地变化。作为实际应用,我们构建了LinkerLLM,一个延迟加载器,它在共驻变体之间别名共享块,实现18-48%的GPU内存节省,并能在单个24 GB消费级GPU上运行多达五个7B参数变体。八种配置中有五种在MMLU、ARC-Challenge、HellaSwag和WinoGrande上保留了至少94%的非共享变体质量;其余三种(Mistral-7B、Gemma-2-2B、Llama-3.2-1B)各自有一个低于阈值的基准(87-91%),我们透明地报告这一点,而不是用单一阈值来门控块共享决策。
英文摘要
Modern LLMs are deployed as families of post-trained variants (base, instruct, chat, code) derived from a shared set of pre-trained weights. We present an empirical study of how post-training transforms weight-space geometry, covering eight configurations across four architecture families (Qwen2.5, Llama-3.1/3.2, Mistral, Gemma-2). We identify a granularity gap: post-training modifies every tensor (zero of 291-339 tensors remain byte-identical, so hash-based deduplication achieves 0% savings), yet preserves block-level structure (mean cosine similarity exceeds 0.99 and relative Frobenius distance stays below 0.13). Post-training therefore acts as a structured perturbation that shifts every parameter while leaving block-level geometry intact. The property is not universal: independently trained specializations (for example, Qwen2.5-Coder) attain cosine similarity around 0.64 with the general base, indicating a disconnected region of weight space. Perturbation magnitude varies systematically with model scale, architecture family, and post-training recipe. As a practical application, we build LinkerLLM, a lazy loader that aliases shareable blocks across co-resident variants, achieving 18-48% GPU memory savings and enabling up to five 7B-parameter variants on a single 24 GB consumer GPU. Five of eight configurations retain at least 94% of the unshared variant's quality on MMLU, ARC-Challenge, HellaSwag, and WinoGrande; the remaining three (Mistral-7B, Gemma-2-2B, Llama-3.2-1B) have one below-threshold benchmark each (87-91%), which we report transparently rather than gate the block-sharing decision on a single threshold.
Comments13 pages, 11 figures. Accepted at the ICML 2026 Workshop on Weight-Space Symmetries: from Foundations to Practical Applications. OpenReview: https://openreview.net/forum?id=jdVfddhT65