发表机构
Colorado School of Mines; Lookia MX(科罗拉多矿业学院; Lookia MX)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出将词汇令牌表示为酉矩阵的方法,无需位置编码即可捕获词序,在文本分类基准上性能优于或相当于词袋基线,且参数效率更高。
AI 中文摘要
我们将词汇令牌表示为酉矩阵,并将每个句子编码为它们的有序乘积。矩阵乘积的非交换性可在不使用位置编码(PEs)的情况下捕获词序。该代数还具备多种能力,包括无需查询、键或值投影的反对称自注意力,以及以更低注意力成本对可变长度文本块进行并行组合。此外,它提供了一个陪集读出层,可紧凑地编码所有真实酉自由度,同时支持通过嵌套群扩展进行持续学习,每次新任务都会扩展算子空间并精确保留先前的表示。在标准文本分类基准上,该方法与词袋基线相当或更优,在IMDB上取得更高准确率,在AG News上表现相当。值得注意的是,这是通过将传统的约30000维词汇空间替换为密集的64参数实值编码实现的,凸显了我们参数化的表达效率。
英文摘要
We represent lexical tokens as unitary matrices and encode each sentence as their ordered product. The noncommutativity of matrix product captures word order without positional encodings (PEs). The same algebra yields several capabilities, including antisymmetric self-attention with no query, key, or value projections, and parallel composition of variable-length text chunks at a reduced attention cost. Furthermore, it provides a canonical-coset readout layer that encodes all true unitary degrees of freedom compactly, while supporting continual learning through nested group extensions that enlarge the operator space with each new task preserving prior representations exactly. Across standard text-classification benchmarks, the method matches or exceeds bag-of-words baselines. Achieving higher accuracy on IMDB and comparable performance on AG News. Notably, this is accomplished by replacing the conventional $\sim$30,000-dimensional vocabulary space with a dense, 64-parameter real-valued encoding, highlighting the expressive efficiency of our parameterization.
Comments11 pages, 3 figures