AI 中文总结
研究针对自回归语言模型序列开头冷启动惩罚问题,提出领域条件位置偏移方法,通过添加单个学习向量减少困惑度,在多模型上效果显著,且具有轻量级、可热切换特点,能改善检索重排等,是短域内评分和校准的有效工具。
AI 中文摘要
自回归语言模型在序列开头准确性最低,此时上下文少,依赖通用预训练先验。研究表明这种冷启动惩罚因领域而异,通过领域条件位置偏移可减少。该偏移是在序列开头位置添加到嵌入激活的单个学习向量,模型权重冻结。它能在约一百个文档上几分钟内训练完成,可在领域间切换且无额外序列状态和可测量延迟开销。在多个模型上,它能将域内困惑度降低达27%,效果在70B模型也持续,一个位置就能捕获大部分益处。匹配的直接逻辑偏差校正最多只能达到7.9%且不改变后续token损失,表明偏移通过模型状态传播而非仅重新校准输出先验。微调的LoRA虽能达到更低困惑度,但使用更多参数。在错误领域控制下,偏移在决策依赖早期域内token时能改善检索重排和领域分类,对信号稍后出现的少样本推理结果不变。位置感知预填充应用对生成任务有帮助,在每个缓存解码步骤应用会导致重复。所以偏移不是最强适配器,而是用于短域内评分和校准的轻量级、热切换工具。
英文摘要
Autoregressive language models are least accurate at the beginning of a sequence, where little context forces reliance on a generic pretraining prior. We show that this cold-start penalty is domain dependent and reduce it with a domain-conditional position offset: a single learned vector added to the embedding activation at the first sequence positions while all model weights remain frozen. The offset trains in minutes on roughly one hundred documents, switches between domains without added sequence state, and has no measurable latency overhead. Across eight Mamba, GPT-NeoX, and Llama models spanning 410M to 8B parameters, it reduces held-out in-domain perplexity by up to 27%; the effect persists at 70B, and one position captures most of the benefit. A matched, converged direct logit-bias correction reaches at most only 7.9% and leaves later-token loss unchanged, showing that the offset propagates through model state rather than merely recalibrating the output prior. A tuned LoRA reaches lower perplexity but uses two to three orders of magnitude more parameters and an active low-rank weight path, while soft prompts add sequence positions. With wrong-domain controls, offsets improve retrieval reranking and domain classification when decisions depend on early in-domain tokens, For the few-shot reasoning whose signal occurs later, the results maintains unchanged. Position-aware prefill application also help generation tasks, whereas naive application at every cached decoding step causes repetition. The offset is therefore not the strongest adapter, but a lightweight, hot switchable tool for short in-domain scoring and calibration.