arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20278cs.LGcs.AI

标记关联结构:用于文本、知识图谱与超图的原生Transformer建模

Labeled Incidence Structures for Native Transformer Modeling of Text, Knowledge Graphs, and Hypergraphs

  • A Carrot, Inc(A Carrot 公司)

机构由 AI 辅助整理,请以论文原文为准。

Mahesh Godavarti

AI总结:

本文提出标记关联结构(LIS),将文本、知识图谱和超图统一编码为$(x_d,s,e)$,使标准Transformer原生处理,并通过算子分解提供角色与关系感知的注意力偏置,同时证明加性编码的局限并分析知识库的容量界。

AI中文摘要:

文本、知识图谱和超图在关系实例中都有扮演不同角色的元素,当数据被扁平化为标记序列时,这种结构信息会丢失。我们引入了标记关联结构(LIS),这是一种统一表示,将每个端点编码为$(x_d, s, e)$:内容$x_d$、角色或槽位$s$,以及该角色出现的关系实例$e$。由于每种数据类型都映射到相同的$(x_d, s, e)$表示而无需扁平化,单个标准Transformer即可原生处理所有类型,结构差异完全由算子承载,而非架构。LIS通过组合槽位算子和实例算子$A(s,e) = R_s R_e$为每个端点分配结构地址。我们刻画了该分解何时赋予每个标记唯一且与路径无关的地址。当满足此条件时,比较端点$j$与端点$i$的自然算子是相对传输$P_{j\ o i} = A_i^{-1} A_j$,这为注意力提供了角色和关系感知的归纳偏置,而无需强加任意序列顺序。形如“位置项加关系项”的加性编码可能丢失依赖于$s$和$e$联合的信息。我们在一个受控示例族中证明了这一点:当旅程算子被近似为仅位置项与仅关系项之和时,该近似无法捕捉位置与关系的组合方式,只能捕捉它们各自的独立效应。我们还分析了持久知识库。与存储位置绑定的标识符使模型对存储顺序敏感,而自由学习的标识符在库规模$M$相对于样本量$n$增长时可能变得更难控制。从内容计算关系实例算子避免了存储顺序问题,并在固定架构和Lipschitz假设下产生独立于$M$的容量界。

英文摘要:

Current Transformer interfaces index tokens by one or more integer coordinates, which determine their addresses inside attention. In RoPE and its multi-axis or hierarchical variants, the resulting address has the form $A(i)=R_1^{i_1}R_2^{i_2}R_3^{i_3}$, where the exponents are integer coordinates assigned after choosing a serialized token layout. When Transformers process new or large collections of data, this addressing scheme can produce unseen offsets or coordinate combinations, push repositories toward retrieve-and-serialize pipelines, and force new entities, records, or repository items to be represented by long token strings or identifier embeddings not seen in training. We introduce labeled incidence structures (LIS), in which each participating token or value is an endpoint with content $x$ and a structural index $i$. The index can include local position, relation role, relation instance, text unit, field, or content-derived identity. The model maps this index to a structural address $A(i)$, so adding new tokens, facts, text units, or repository items applies the same learned address rule to structural and content coordinates rather than requiring larger integer coordinates, unseen coordinate combinations, or new identifier embeddings. Attention scores endpoints $i,j$ using $q_i^\top P_{j\to i}k_j$, where journey consistency forces $P_{j\to i}=A(i)^{-1}A(j)$. When $i$ has several coordinates, such as position, role, and instance, coordinate independence is equivalent to factoring $A(i)$ into one address factor per coordinate. This recovers RoPE, RoPE-2D, and HiRoPE as special cases. This allows knowledge-graph (KG) roles, fact instances, and text units to enter the attention score directly. In controlled shallow diagnostics, the LIS address interface is implemented inside ordinary Transformer attention and yields promising results across text, KG, and $n$-ary settings.

↑