拓扑引导
Topological Steering
浏览论文内容
中文总结 AI 辅助
针对现有大语言模型行为控制方法易受局部扰动影响的问题,提出基于拓扑数据分析的Topological Steering框架,通过激活空间的拓扑表示实现更鲁棒的模型行为控制,且在多模型家族及规模上均有效。
中文摘要 AI 辅助
随着大语言模型(LLMs)的快速兴起,控制模型的不良行为变得愈发重要。现有的行为控制方法通常直接干预激活空间或特征空间,但这类方法对异常值、分布偏移、噪声及其他局部扰动较为敏感。受拓扑数据分析(TDA)(一种捕捉全局而非纯局部结构的方法)启发,我们提出了Topological Steering(拓扑引导)这一新框架,通过激活空间的拓扑表示来引导大语言模型的行为。该方法利用持续同调图,将基于激活的引导与拓扑数据分析相结合,实现了更鲁棒的行为控制。我们证明,Topological Steering可在多个模型家族及不同模型规模下一致地修改大语言模型的行为。
英文摘要
With the rapid rise of large language models (LLMs), controlling undesirable model behaviors has become increasingly important. Existing behavioral control methods typically intervene directly in activation or feature space, but such approaches can be sensitive to outliers, distributional shifts, noise, and other local perturbations. Motivated by Topological Data Analysis (TDA), which captures global rather than purely local structure, we propose Topological Steering, a new framework for steering LLM behavior through the topological representation of activation spaces. Using persistence diagrams, our method connects activation-based steering with TDA and enables more robust behavioral control. We show that Topological Steering consistently modifies LLM behavior across multiple model families and model sizes.
发表机构
- National University of Singapore (NUS)(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。