规划还是即兴?对开放模型和开放跨层转码器的诗歌规划位点进行压力测试
Planning or Improvisation? Stress-Testing the Poetry Planning Site on Open Models and Open Cross-Layer Transcoders
浏览论文内容
中文总结 AI 辅助
对开放模型和跨层转码器进行压力测试,发现押韵规划的位置特异性仅在生成相邻处成立,新行符驻留规划未被恢复,表明因果位点紧邻生成,构成边界条件而非反驳。
中文摘要 AI 辅助
Lindsey等人(2025)报告称,Claude 3.5 Haiku会规划押韵:候选押韵词的特征在写出一行之前的新行符上被激活,并且抑制-注入干预仅在该位置应用时才能重定向该行(他们的图13)。我们测试了这种泛化在多大程度上适用于跨越四个开放模型(参数规模从0.6B到2.6B)和六个开放跨层转码器(CLT)的七个单元上,使用一块消费级GPU,将该论断分解为位置特异性(C1)、新行符位点同一性(C2)和新行符驻留规划(C3)。这是一次压力测试而非忠实复现:这些CLT无法获得归因图,因此特征是从解码器向量自下而上发现的。C1在所有单元以及444个提示-注入对中具有可检测效应的全部247对中均成立,但有效位置是紧邻生成的最终提示词元,且只有两个单元达到了行为上有意义的概率。C2和C3未被任何探针恢复:对所有活跃特征进行普查发现,在新行符处没有押韵预期的富集,并且在模型组合整行时对新行符进行引导,经过36次运行和8,640个采样行,揭示了原因。该干预虽然强烈但仅持续一个词元,使得注入的词在720个样本中的703个中成为组合行的第一个词,而押韵词则在六个词之后未被触及。最终测试完全弃用转码器:在每一层修补新行符的整个残差,使用第三行以不同押韵结尾的最小对诗歌,在1,260个组合行中移动了押韵11次,而基线为4次,设计解析率为1.4%。我们将此解读为边界条件而非反驳:在此规模和使用这些转码器的情况下,因果位点紧邻生成。我们复现了图13的形状,而非其机制。代码和数据已公开(代码:此HTTP URL)。
英文摘要
Lindsey et al. (2025) report that Claude 3.5 Haiku plans rhymes: features for candidate rhyme words are active on the newline before a line is written, and a suppress-and-inject intervention redirects the line only when applied there (their Figure 13). We test how far this generalizes on seven cells crossing four open models (0.6B to 2.6B parameters) with six open cross-layer transcoders (CLTs), on one consumer GPU, decomposing the claim into position specificity (C1), newline site identity (C2), and a newline-resident plan (C3). This is a stress test rather than a faithful reproduction: attribution graphs are unavailable for these CLTs, so features are found bottom-up from decoder vectors. C1 generalizes, in every cell and in all 247 of 444 prompt-by-inject pairs with a detectable effect, but the effective position is the final prompt token, adjacent to emission, and only two cells reach behaviorally meaningful probabilities. C2 and C3 are not recovered by any probe: a census of every active feature finds no rhyme-anticipating enrichment at the newline, and steering the newline while the model composes the whole line, over 36 runs and 8,640 sampled lines, shows why. That intervention is strong but one token long, making the injected word the first word of the composed line in 703 of 720 samples and leaving the rhyme six words later untouched. A final test drops the transcoder entirely: patching the newline's whole residual, at every layer, from a minimal-pair poem whose third line ends on a different rhyme moves the rhyme in 11 of 1,260 composed lines against 4 at baseline, with a design resolving 1.4%. We read this as a boundary condition rather than a refutation: at this scale and with these transcoders, the causal site is emission-adjacent. We reproduce Figure 13's shape, not its mechanism. Code and data are public (code: github.com/PCfVW/poetry-planning-site).
发表机构
- Cosmic AI(宇宙人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。