AI 中文总结
该研究发现激活探针所需合成数据量由覆盖度而非难度决定,指令类探针需更多样本,高风险类探针样本利用率高,且公开了相关数据集。
AI 中文摘要
监测部署型语言模型的激活探针基于合成对话训练,所需样本量尚不明确。本研究针对高风险场景、对人有害的回复、不遵循用户指令的回复三类监测概念,在14个保留的评估分布、4种探针模型上,追踪了10至590个合成样本的学习曲线,同时改变生成器大语言模型(LLM)和提示词的细节。研究发现,所需样本量由监测对象决定:在Gemma-3-27B-IT上,高风险和有害类探针在80个样本时就达到 plateau 的99%左右,指令类探针则需要数倍于此的样本量,且该规律在3种更小的探针模型及真实样本(来自开发集)上均成立。现有研究建议将生成预算投入到广度(更多数据种类)而非深度(更多同种类数据)。本研究将某概念所需的深度定义为拟合曲线的半增益规模,即达到一半增益时所需的样本数。概念和分布可解释42%-45%的方差,生成器、探针模型和提示词细节的贡献不足10%。决定半增益规模的是覆盖度而非同种类难度:某分布达到饱和所需的同种类样本数。所有三类概念下,每种类型(对应一个评估分布)的同种类合成样本的半增益规模中位数均为7-11。差异在于某类型样本向该概念其他类型的迁移程度:高风险场景下几乎完全迁移,有害类迁移程度较低,指令类迁移程度最低,这解释了生成样本和真实样本上不同概念间的大部分差距。因此,广度的价值因概念而异:指令类下,因无类型可覆盖其他类型,需多种类型;高风险场景下,因一种类型可覆盖其余类型,多种类型近乎冗余。本研究公开了评估套件、开发集和生成集。
英文摘要
Activation probes that monitor deployed language models are trained on synthetic conversations, and how many a probe needs is open. We trace learning curves over 10-590 synthetic samples for three monitoring concepts, high-stakes situations, replies harmful to a person, and replies that do not follow the user's instruction, on fourteen held-out evaluation distributions and four probe models, varying the generator LLM and the prompt's detail. The need is set by what is monitored: probes for high-stakes and harmful are within a few hundredths of their plateau from 80 samples on Gemma-3-27B-IT, instruction probes need several times as many, and the ordering holds on three smaller probe models and on real samples (from dev set). Prior work advises spending a generation budget on breadth, more kinds of data, over depth, more of each kind. We read the depth a concept needs as the half-gain size of a fitted curve, the number of samples at which half the gain is in hand. Concept and distribution account for 42-45% of its variance, the generator, probe model, and prompt detail for under 10%. What sets the value of the half-gain size is coverage, not per-kind difficulty: the number of samples of its own kind a distribution needs to saturate. Every kind, one per evaluation distribution, has a median half-gain size of 7-11 own-kind synthetic samples under all three concepts alike. What differs is how far samples of one kind transfer to the concept's other kinds, almost fully under high-stakes, less under harmful, and least under instruction, which accounts for most of the gap between concepts on generated and real samples. Breadth therefore pays differently by concept: many kinds are necessary under instruction, where no kind covers another, and nearly redundant under high-stakes, where one kind covers the rest. We release the evaluation suites, dev sets, and generated sets.
Comments24 pages, 8 figures, 11 tables. Code and data: https://github.com/Ankush7890/sythentic_data_needs