arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越线性表示假设:文本到图像模型中的非线性激活引导

Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models

Muhammad Atif Butt, Paweł Skierś, Joost Van De Weijer, Kamil Deja

arXiv 2610.06945首次发表:更新:

发表机构

Computer Vision Center; Universitat Autònoma de Barcelona; Warsaw University of Technology; IDEAS Research Institute(计算机视觉中心; 巴塞罗那自治大学; 华沙理工大学; IDEAS研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到图像模型中概念过渡的非线性特性,提出KANSteer方法,利用Kolmogorov-Arnold网络建模曲线激活轨迹,实现比线性引导更平滑的中间属性遍历。

AI 中文摘要

机制可解释性通常依赖于线性表示假设(LRH),该假设认为高层概念在激活空间中被编码为线性方向。然而,自然的视觉概念并不一定需要线性的视觉过渡:在晴朗和暴风雨之间存在着一个中间天气状态,例如带有几朵白云的天空,而不仅仅是较弱的暴风雨;在毛毛虫和蝴蝶之间,其进展也不是一只翅膀不断生长的毛毛虫。这引发了一个问题:这种真正的中间状态是否也被模型非线性地表示。确实,当我们直接提示文本到图像模型生成中间属性时,其激活很少落在连接端点的直线方向上。因此,我们提出了KANSteer,将概念遍历建模为一条穿过其中间状态的曲线。为了寻求既简单又可解释的表示,我们建议使用Kolmogorov-Arnold网络(KANs),它提供一个一维坐标,其学习到的函数定义了轨迹。这使得引导方向在概念上可以变化,同时保持可解释的表示。在多个概念和文本到图像扩散Transformer上,我们发现它们的激活轨迹显著偏离直线,并且KANSteer比线性引导提供了更紧密的拟合和更平滑的中间属性遍历。

英文摘要

Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear directions in activation space. Yet a natural visual concept does not necessarily require a linear visual transition: between sunny and stormy lies an intermediate weather state such as a sky with a few white clouds, not simply a weaker storm; between a caterpillar and a butterfly, the progression is not a caterpillar with continuously growing wings. This raises the question of whether such true intermediate states are also represented nonlinearly by the model. Indeed, when we prompt text-to-image models directly for intermediate attributes, their activations rarely fall along the straight direction connecting the endpoints. Therefore, we propose KANSteer, which models concept traversal as a curve passing through its intermediate states. Seeking a representation that is both simple and interpretable, we propose to use Kolmogorov-Arnold Networks (KANs), which provide a one-dimensional coordinate whose learned functions define the trajectory. This allows the steering direction to vary along the concept while preserving an interpretable representation. Across several concepts and text-to-image diffusion transformers, we find that their activation trajectories substantially deviate from straight lines, and that KANSteer provide a closer fit and smoother traversal of intermediate attributes than linear steering.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑