用于结构化多轴组合的旅程算子
Journey Operators for Structured Multi-Axis Composition
浏览论文内容
中文总结 AI 辅助
该研究提出旅程算子框架,建模多轴数据结构,统一RoPE等方法,设计JoFormer模型,实验验证其归纳偏置在视觉、语言等任务中有效。
中文摘要 AI 辅助
许多类型的数据在一个或多个轴上具有结构:句子中的单词、图像中的像素、树中的节点、音频中的帧或3D体积中的单元。沿某一轴,顺序很重要:“狗咬人”与“人咬狗”不同。然而,在独立轴之间,组合或移动不应依赖于轴的顺序:在图像中,先向右再向下组合应与先向下再向右组合结果相同,且先向右再向下移动应与先向下再向右移动描述相同的相对位置。我们开发了一种用于建模此类多轴结构的框架。每个数据项携带其内容以及每个轴的小型变换。连接两个位置的路径定义了一段旅程;旅程算子是沿该路径的各轴变换的乘积,既控制数据沿路径的组合方式,也控制相对位置的描述方式。当变换固定时,我们的框架会恢复旋转位置嵌入(Rotary Position Embedding, RoPE)及其多维变体。当变换依赖于数据时,模型会获得内容自适应的位置归纳偏置。我们明确证明了这些路径何时是良定义的:仅当轴变换可交换时,跨轴的组合和移动才是路径无关的。我们还证明,在给定的环面框架对称性、上同调、双线性性和保范假设下,所得的成对评分规则必须采用分块旋转的形式,解释了为何类RoPE方法会自然出现。最后,我们利用该理论设计了用于值聚合的模型JoFormer,并将其与注意力机制和状态空间模型(State-Space Models, SSMs)关联起来。在视觉、语言和长度泛化方面的初步实验表明,这些归纳偏置在实践中可产生可观测的效果。
英文摘要
Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. Along one axis, order matters: "the dog bit the man" is different from "the man bit the dog." Across independent axes, however, neither composition nor movement should depend on the order of axes: in an image, composing right then down should give the same result as composing down then right, and moving right then down should describe the same relative position as moving down then right. We develop a framework for modeling this kind of multi-axis structure. Each data item carries its content together with a small transformation for each axis. A path connecting two positions defines a journey; the journey operator is the product of per-axis transformations along that path, governing both how data composes along the path and how relative position is described. When the transformations are fixed, our framework recovers Rotary Position Embedding (RoPE) and its multi-dimensional variants. When they depend on the data, the model gains a content-adaptive positional inductive bias. We show exactly when these paths are well-defined: both composition and movement across axes are path-independent precisely when the axis transformations commute. We also prove that, under the stated toral-frame symmetry, cocycle, bilinearity, and norm-preservation assumptions, the resulting pairwise scoring rule must take the form of block-wise rotations, explaining why RoPE-like methods arise naturally. Finally, we use this theory to design JoFormer, a model for value aggregation, and relate it to attention and state-space models (SSMs). Initial experiments across vision, language, and length generalization suggest that these inductive biases can have observable consequences in practice.
发表机构
- A Carrot, Inc(A Carrot公司)
机构由 AI 辅助整理,请以论文原文为准。