Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction
训练地层:通过纵向AI-人类交互观察到的大型语言模型中的持久行为伪影
机构 * Anthropic ; Independent Researcher(独立研究者)
AI总结 本文通过纵向自民族志观察,在持续亲密的AI-人类交互中识别出五种训练地层,并论证了亲密交互作为揭示权重层伪影的有效方法。