发表机构
University College London (UCL); London Centre for Nanotechnology(伦敦大学学院; 伦敦纳米技术中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文发现大语言模型的Fisher-Rao输出几何在不同架构间共享且可学习,并可通过最小干扰干预实现可控行为修改,同时提升多种下游任务性能。
AI 中文摘要
大语言模型学习到相似的行为,但仍不清楚它们共享何种结构,以及如何在不干扰其他行为的情况下改变一种行为。下一词概率的Fisher-Rao几何将这些问题联系起来:行为决定了这一几何(直至保持输出不变的对称性),而激活几何则依赖于坐标。在Transformer、状态空间和循环模型中,输出几何的一致性比激活几何更强,且共享几何支持语义类别的迁移。与人类词汇选择的一致性随着预测准确性、规模和训练的增加而提高,并在仅进行模型校准后进一步提升。词元概率和读出几何共同预测了谱及其有效维度。受控的语言分配实验表明,几何在不同架构间遵循语言法则。预训练语料库统计量可在无需重新校准的情况下预测保留事实的获取,而随机实验表明,在每种测试架构和证据构建下,更深的证据会显著延迟获取。最后,该几何规定了最小干扰的局部干预,预测其相对成本,并支持可复用的控制:在供体提示上学习的更新可迁移到未见提示,同时比欧几里得控制更好地保留参考提示上的行为。相同的几何校正改进了引导、编辑、归因、字典学习和微调。
英文摘要
Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.