内省微调(IFT):训练小型语言模型进行内省
Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
浏览论文内容
中文总结 AI 辅助
研究小型语言模型内省能力,提出句子定位和强度比较评估范式,发现小模型有内省能力且随规模提升,引入内省微调(IFT),能提高模型内省能力,还表明内省能力可训练,对AI透明度和对齐有意义。
中文摘要 AI 辅助
我们通过激活引导的视角来研究小型语言模型能否检测并报告其自身内部激活的扰动。具体做法是将概念向量注入模型的残差流,并衡量模型是否能准确报告该扰动。研究发现先前工作中的二元检测范式在小型模型中存在混淆,因此提出了句子定位和强度比较这两种无混淆评估范式。通过对六个模型的评估发现,小至2B参数的模型内省能力可靠且高于随机水平,内省能力一般随规模增加。还引入了内省微调(IFT),它能提高模型的内省能力,同时对标准能力基准的影响可忽略不计。研究结果表明内省能力并非仅由规模决定,可直接训练,这对人工智能的透明度和对齐有影响。
英文摘要
Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens of activation steering: injecting concept vectors into a model's residual stream and measuring whether the model can accurately report on the perturbation. We first show that the binary detection paradigm used in prior work -- prompting the model to answer Yes'' or No'' to whether it detects an injected thought -- is confounded in small models, as steering biases the model toward affirmative responses regardless of the question content. We therefore propose two confound-free evaluation paradigms: sentence localization (identifying which of $N$ sentences was perturbed, chance $= 1/N$) and strength comparison (identifying which of two sentences received a stronger injection, chance $= 50\%$). Evaluating across six models from two families (Llama-3.2 and Gemma-4), we find that models as small as 2B parameters introspect reliably well above chance, and that introspective ability generally increases with scale. Llama-1B, however, performs at or below chance. We then introduce \emph{Introspection Fine-Tuning} (IFT): supervised fine-tuning on sentence-localization examples constructed from the model's own perturbed forward passes. IFT raises Llama-1B sentence-localization accuracy from $9.6\%$ to $60.6\%$ (a $6\times$ improvement), with gains generalizing zero-shot to the held-out strength-comparison task ($30.2\% \to 52.2\%$). IFT also improves introspection for 3B and 8B models, while inducing negligible degradation on standard capability benchmarks. Our results suggest that introspective ability is not fixed by scale alone: it can be directly trained, and doing so unlocks latent self-monitoring capacity with implications for AI transparency and alignment. Our code is \href{https://anonymous.4open.science/r/IFT-introspection-2092/README.md}{here}.
发表机构
- Harvard College(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。