AI 中文总结
本研究提出SemanticFold潜在序列压缩方法,发现其在不同模型上对语言建模、可解码性和推理的影响非单调且方向各异,表明压缩保持无单一标量指标。
AI 中文摘要
我们研究了提示前缀的潜在序列压缩是否保留大型语言模型在推理过程中依赖的能力。我们引入了SemanticFold,一种在学习到的边界处折叠前缀隐藏状态的压缩方案,并在五个模型规模上进行了评估:Qwen3-1.7B、Qwen3-8B、SmolLM2-1.7B、Pythia-1.4B和Pythia-6.9B。我们使用固定目标协议:冻结的前缀以原生方式或压缩方式执行,两个分支教师强制相同的续接标记。该设计排除了目标选择对似然变化的解释。我们考察了五个端点家族:固定目标负对数似然、有限标签推理准确率、线性探针可访问性、开放式生成以及系统级内存和延迟。我们发现压缩使这些端点非单调变化,且它们不共享单一压缩阈值。在Qwen3-1.7B上,压缩比R=1.7时,在10000次抽取的配对自助法下,压缩减去原生的平均NLL降低0.135。在SmolLM2上,R=1.2时,平均变化比原生高0.013。在两个Pythia检查点上,NLL实际上不变。将序列缩短与学习到的残差变换分离的NLL分解表明,Qwen的有利似然主要归因于残差适应而非仅缩短。仅MLP(应用变换而不缩短)比完整SemanticFold的NLL低0.082。线性探针准确率和宏AUC在各条件下的绝对值变化小于0.03,置信区间跨越零。我们得出结论,潜在压缩下的保持没有单一标量证书:语言模型拟合、可解码性和推理行为回答不同的问题,在同一压缩操作下可能朝不同方向移动。
英文摘要
We study whether latent sequence compression of prompt prefixes preserves the capabilities that large language models rely on during inference. We introduce SemanticFold, a compression scheme that folds prefix hidden states at learned boundaries, and evaluate it across five model scales: Qwen3-1.7B, Qwen3-8B, SmolLM2-1.7B, Pythia-1.4B, and Pythia-6.9B. We use a fixed-target protocol: a frozen prefix is executed natively or compressed, and both arms teacher-force identical continuation tokens. This design rules out target-selection explanations for likelihood changes. We examine five endpoint families: fixed-target negative log-likelihood, finite-label reasoning accuracy, linear probe accessibility, open-ended generation, and systems-level memory and latency. We find that compression moves these endpoints non-monotonically and that they do not share a single compression threshold. On Qwen3-1.7B at compression ratio R=1.7, compressed-minus-native mean NLL decreases by 0.135 under paired bootstrap with 10000 draws. On SmolLM2 at R=1.2, the mean change is 0.013 higher than native. On both Pythia checkpoints, NLL is effectively unchanged. An NLL decomposition separating sequence shortening from the learned residual transform shows that the favorable Qwen likelihood is attributable primarily to residual adaptation rather than to shortening alone. MLP-only, which applies the transform without shortening, achieves 0.082 lower NLL than Full SemanticFold. Linear probe accuracy and macro AUC change by less than 0.03 in absolute value across conditions, with confidence intervals crossing zero. We conclude that preservation under latent compression has no single scalar certificate: language-model fit, decodability, and reasoning behavior answer different questions and can move in different directions under the same compression operation.