发表机构
Tel Aviv University; Google Research; Columbia University(特拉维夫大学; 谷歌研究院; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究如何通过简单API级控制调整语言模型,开发黑箱方法学习对数偏差向量,无需修改模型权重或梯度,该方法在数学和推理基准上优于基础模型,是一种轻量级的语言模型适应机制。
AI 中文摘要
许多组织旨在调整语言模型以供内部使用,以提高特定领域任务的性能并解决敏感数据的隐私问题。然而,这种调整并非易事:通常需要对开源模型进行具有操作挑战性的微调或临时提示优化。我们研究了一种基于简单API级控制的最小替代方案:允许用户使用用户定义的向量对模型的对数进行偏差调整。我们开发了一种黑箱方法,用于学习在每个解码步骤添加的单个上下文无关的对数偏差向量,而无需修改模型权重或梯度。从KL正则化强化学习目标出发,我们刻画了这种固定对数偏差向量何时可以近似最优前缀依赖校正,并从展开、奖励和令牌概率中导出封闭形式的逆倾向估计器。从经验上看,这种简单的解码时间干预在数学和推理基准上优于基础模型,同时使用的可训练参数比传统微调少得多。我们的结果表明,学习到的对数偏差是一种在最小访问要求下调整语言模型的轻量级机制。
英文摘要
Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address privacy concerns around sensitive data. However, such adaptation remains non-trivial: it often requires operationally challenging fine-tuning of open-source models or ad hoc prompt optimization. We study a minimal alternative based on a simple API-level control: allowing users to bias the model's logits with a user-defined vector. We develop a black-box method for learning a single context-independent logit-bias vector, added at every decoding step, without modifying model weights or requiring gradients. Starting from a KL-regularized reinforcement learning (RL) objective, we characterize when such a fixed logit-bias vector can approximate the optimal prefix-dependent correction and derive a closed-form inverse-propensity estimator from rollouts, rewards, and token probabilities. Empirically, this simple decoding-time intervention improves over base models on mathematical and reasoning benchmarks while using far fewer trainable parameters than conventional fine-tuning. Our results suggest that learned logit bias is a lightweight mechanism for adapting language models under minimal access requirements.
Comments36 pages, 4 figures, 6 tables. Appendix included. Code available at https://github.com/Ofek-Israeli/logit_bias_lm_adaptation