发表机构
Institute of Engineering, Thapathali Campus; Tribhuvan University(工程学院塔帕塔利校区; 特里布文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统评估了量化大语言模型中激活操控的剂量-反应关系,发现情感操控在量化下保持有效,而推理长度操控呈现不对称失败模式,并揭示了方法论陷阱与压缩对操控的潜在主导效应。
AI 中文摘要
推理时激活操控能够在无需修改参数的情况下实现对大型语言模型的行为控制,而训练后量化则降低了部署时的内存和计算成本。尽管这两种技术在实践中日益融合,它们之间的相互作用仍未得到表征。我们系统地研究了在仅权重量化(INT8和NF4)下,跨四个开放权重7-9B模型和两个行为目标(评判情感和无评判推理长度)的激活操控。使用一种等效应框架,在匹配行为效应下比较能力代价,我们发现情感操控在量化下完好无损。在用统一的v2.3.1重评分修正GSM8K解析器伪影后,合并的INT8对比度为-0.010(90%置信区间[-0.026, +0.007]),在预注册的三标签规则下描述性地等价,而NF4仍为不确定,为-0.017([-0.067, +0.033])。相比之下,推理长度表现出令人惊讶的不对称剂量-反应:延长是渐进的,但终止于上限失控和崩溃,而缩短是阶跃函数,在非连续失败前仅缩短12-30%(因模型而异)。我们揭示了一个方法论陷阱:朴素的等效应阶梯在目标有下限时锚定在崩溃底板上,我们引入了一种删失构造,恢复了可解释的交叉点。我们还量化了Mistral-NF4的显著基线能力偏移(alpha=0时GSM8K从0.545降至0.365),表明压缩可以主导操控干预。尽管如此,操控向量与其FP16对应物保持高度共线(INT8余弦相似度为0.989-0.998,NF4为0.945-0.990),确认了行为方向在量化下幸存,即使代价结构不复存在。所有代码和数据均已发布。
英文摘要
Inference-time activation steering enables behavioral control of large language models without parameter modification, while post-training quantization reduces memory and compute costs for deployment. Despite their growing convergence in practice, the interaction between these two techniques remains uncharacterized. We systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length. Using an iso-effect framework that compares capability costs at matched behavioral effect, we find that sentiment steering survives quantization intact. After correcting a GSM8K parser artifact with a uniform v2.3.1 rescore, the pooled INT8 contrast is -0.010 (90% CI [-0.026, +0.007]), descriptively Equivalent under the preregistered three-label rule, while NF4 remains Inconclusive at -0.017 ([-0.067, +0.033]). In contrast, reasoning length exhibits a surprising asymmetric dose-response: lengthening is graded but terminates in cap-runaway and collapse, while shortening is a step function with only 12-30% shortening (model-dependent) before discontinuous failure. We expose a methodological pitfall: the naive iso-effect ladder anchors on the collapse floor for floor-bounded targets, and we introduce a censored construction that restores interpretable crossings. We also quantify a substantial baseline capability shift for Mistral-NF4 (0.545 to 0.365 GSM8K at alpha=0), demonstrating that compression can dominate the steering intervention. Despite this, steering vectors remain highly collinear with their FP16 siblings (cosine similarity 0.989-0.998 for INT8, 0.945-0.990 for NF4), confirming that the behavioral direction survives quantization even when the cost structure does not. All code and data are released.
Comments17 pages, 5 figures