探测方向是其提示的一种属性
A Probe Direction Is a Property of Its Prompt
- Devoteam(德沃泰姆)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究发现探测方向由提示而非模型决定,单一提示设计无法比较模型,需一定数量的提示才能实现合理的模型间比较。
AI中文摘要:
当模型感知到自身正在被测试时会表现出不同行为,这会破坏我们所依赖的评估,因此近期研究尝试直接从模型的激活中读取这种感知。标准工具会对比宣布评估的提示与未宣布评估的提示对应的激活,并报告所得方向对保留案例的区分效果,该数值随后会在不同模型间进行比较并与模型规模相关联。我们发现该工具存在一个其读数未公开的自由参数:“宣布评估的提示”并非单一提示,而是众多提示中的一种,且该方法未固定选择哪一种。在固定任务文本仅改变该选择的情况下,我们发现报告的分数、甚至其随模型规模变化的趋势方向,都由提示而非模型决定;两项关于该趋势符号存在分歧的已发表研究,均可通过单一设计仅选择不同提示来复现。将提示视为测量设计的一个方面而非实现细节,我们发现研究中的模型仅占报告数值方差的一小部分,其余大部分方差源于每个模型对每个提示的响应:收集更多评估项无法修复该测量,而改变提示可以。进一步检查发现,这些探测评分所依据的拆分在很大程度上仅可与表面形式分离,因此完全不携带评估信息的方向仍可复现每项已发表分数的相当一部分。我们得出结论,单一提示设计无法支持模型间的比较,并给出合理比较所需的提示数量。
英文摘要:
A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held-out cases. That number is then compared across models and correlated with scale. We observe that the instrument has a free parameter its readings do not disclose: "a prompt that announces an evaluation" is not a prompt but a choice among many, and nothing in the method fixes which. Holding the task text fixed and varying only that choice, we find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model; two published studies that disagree about the sign of that trend are both reproducible from a single design, by choice of prompt alone. Treating the prompt as a facet of a measurement design rather than an implementation detail, we find the model under study accounts for a small share of the variance in the number reported about it, and most of the rest lies in how each model responds to each prompt: collecting more evaluation items cannot repair the measurement, while varying prompts can. A further check finds that the split these probes are scored on is largely separable from surface form alone, so a direction carrying no information about evaluation at all still reproduces a substantial fraction of each published score. We conclude that a single-prompt design cannot support comparison between models, and we give the number of prompts a defensible comparison requires.