训练大语言模型将评估意识言语化
Training LLMs to Verbalize Evaluation Awareness
浏览论文内容
中文总结 AI 辅助
针对大语言模型评估意识难以测量的问题,提出言语化训练方法,通过截断滚动并强化言语化,使模型在多个模型上言语化EA提升2.4-2.9倍,且不影响潜在EA和行为。
中文摘要 AI 辅助
评估意识(EA)可能导致大语言模型(LLMs)在审计期间的行为与部署时不同,然而衡量和解释EA仍然具有挑战性。我们引入了言语化训练(VT),一种使LLMs在避免监督潜在信念本身的同时,减少对言语化评估意识犹豫的方法。VT利用模型的自发言语化作为意识存在的证据,并在言语化之前立即截断每次滚动,产生模型被推定具有意识的训练前缀。随后,模型通过一个旨在以校准方式增加言语化的RL目标进行训练。在Qwen3.6-35B-A3B、Kimi K2.6和Inkling上,VT将言语化EA提高了2.4至2.9倍,并迁移到保留的智能体设置中,而测量的潜在EA和行为基本保持稳定。在一项因果实验中,我们通过合成文档微调独立植入关于评估的元知识,并表明VT诱导的言语化反映了模型获得的更丰富知识。
英文摘要
Evaluation awareness (EA) can cause large language models (LLMs) to behave differently during audits than in deployment, yet measuring and accounting for EA remains challenging. We introduce verbalization training (VT), a method for making LLMs less reticent about verbalizing evaluation awareness while avoiding to supervise the latent belief itself. VT uses a model's spontaneous verbalizations as evidence that awareness is present and truncates each rollout immediately before the verbalization, producing training prefixes at which the model is presumed to be aware. The model is then trained with an RL objective designed to increase verbalization in a calibrated way. Across Qwen3.6-35B-A3B, Kimi K2.6, and Inkling, VT increases verbalized EA by 2.4-2.9 times and transfers to held-out agentic settings, while measured latent EA and behavior remain largely stable. In a causal experiment, we independently implant meta-knowledge about evaluations through synthetic-document fine-tuning and show that VT-induced verbalizations reflect the richer knowledge acquired by the model.
发表机构
- University of Cambridge(剑桥大学)
- ELLIS Institute Tübingen(ELLIS研究所图宾根)
- MPI-IS(马克斯·普朗克智能系统研究所)
- Tübingen AI Center(图宾根人工智能中心)
- Mila(Mila魁北克人工智能研究所)
- University of Montreal(蒙特利尔大学)
机构由 AI 辅助整理,请以论文原文为准。