AI 中文总结
本文发现基础模型可通过固定起始标记线索(如“。\n\n好的”)显著提升推理性能,甚至媲美强化学习模型,并证明其效果源于训练数据中的关联,且影响安全行为。
AI 中文摘要
在本文中,我们研究了训练数据如何在基础模型响应的起始标记与后续推理行为之间建立关联。首先,我们证明固定特定的起始标记线索可以使基础模型在数学和编程任务上的性能与经过强化学习(RL)训练的模型相当。例如,线索“。\n\n好的”将Olmo-3-7B在MATH-500上的pass@1准确率从42%提升至78%,而“好吧,”将Qwen3-14B的准确率从72%提升至87%。其次,强化学习使这些线索更有可能出现,而固定这些线索可以恢复其在基础模型上的大部分性能提升。第三,我们将标记线索的推理效果追溯到训练数据。我们执行因果数据干预,将任意单词(如“鸡”)转变为有效的推理线索,或消除现有线索的效果。类似的编辑使得提示指令“思考鸭鸭鹅”在引发推理方面与“逐步思考”同样有效。我们还发现,不同线索引发的隐藏状态表示与训练集中的不同文档类型相关。最后,我们通过语言模型安全性的案例研究扩展了我们对标记线索的研究,发现不同线索会引发不同的拒绝和顺从行为,这些行为对应于不同类型的训练数据。
英文摘要
In this paper, we study how training data creates associations between the tokens at the start of a base model's response and the reasoning behavior that follows. First, we demonstrate that fixing particular starting token cues makes a base model's performance competitive with that of its reinforcement learning (RL)-trained counterparts on math and coding. For instance, the cue ".\n\nOkay" raises Olmo-3-7B's MATH-500 pass@1 accuracy from 42% to 78%, while "Alright," raises Qwen3-14B's from 72% to 87%. Second, RL makes these cues more likely, while fixing them recovers much of its performance gain over the base model. Third, we trace the reasoning effects of token cues to the training data. We perform causal data interventions to turn an arbitrary word, such as "chicken", into an effective reasoning cue, or remove an existing cue's effect. A similar edit makes the prompt instruction "Think duck duck goose" as effective as "Think step by step" at eliciting reasoning. We also find that the hidden state representations induced by different cues correlate with different document types from the training set. Finally, we extend our study of token cues with a case study in language model safety, finding that different cues elicit distinct refusal and compliance behaviors that correspond to different types of training data.
CommentsProject page: https://www.sophielwang.com/cues Code: https://github.com/sophicle/cues